Big language model construction method and system for identifying bidding and tendering data

By constructing a large language model, the efficiency and accuracy issues of traditional bidding data recognition technology are solved, enabling efficient processing of multimodal data and intelligent analysis of legal clauses, improving the legality and compliance of bidding data, and providing more comprehensive information support.

CN121168652APending Publication Date: 2025-12-19BEIJING AODETA DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511328976.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Traditional bidding data recognition technologies struggle to handle complex multimodal data, lack professional knowledge about the bidding field, leading to misunderstandings and misjudgments. Furthermore, the lack of domain-adaptive fine-tuning methods affects the assessment of legality and compliance.

Method used

A large language model is constructed by acquiring multimodal bidding data, performing structured processing and format standardization, building a legal terminology knowledge graph, determining the weights and semantic relationships of legal terms, cleaning redundant information, constructing a legal clause network graph, detecting conflicts and generating risk reports, and realizing incremental learning and fine-tuning of the model.

Benefits of technology

It improves the efficiency and accuracy of bidding data identification, enhances regulatory capabilities, enables insights into market supply and demand and industry trends, helps enterprises flexibly respond to legal provisions, and improves the accuracy and comprehensiveness of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168652A_ABST
    Figure CN121168652A_ABST
Patent Text Reader

Abstract

The invention provides a big language model construction method and system for identifying bidding data, equipment and a medium. The method comprises the following steps: constructing a basic term dictionary according to a legal term version, determining a semantic association relationship between legal terms in combination with a restart random walk algorithm, and constructing a legal term knowledge graph; determining the level of legal terms in the basic term dictionary, and determining the weight coefficient of the legal terms through a preset time attenuation factor; obtaining a basic term dictionary and a legal term domain dictionary; determining a reference relationship between the discriminants based on the discriminant text, and constructing a directed acyclic legal clause network diagram; determining the effectiveness value of each processed legal clause as the value of each node; determining a potency propagation relationship between legal terms, obtaining an edge weight, and obtaining a weighted term network; generating a bidirectional weighted correlation graph; and extracting the weight of each legal term in the optimized basic dictionary by utilizing a bidirectional weighted correlation graph, and performing incremental learning on a demand layer of a large language model by utilizing model fine tuning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, in particular to a large language model construction method, system, device and medium for identifying bidding data. BACKGROUND

[0002] With the rapid popularization of mobile Internet, information technology and bidding and tendering business are deeply integrated, and bidding and tendering data show explosive growth and massive aggregation, which has a very significant impact on bidding work. These data cover various modalities such as text, table, image, and involve content such as bidding announcement, tender offer, technical scheme, commercial offer, expert review opinion and scoring record.

[0003] Traditional bidding document review relies on manual or rule-based natural language processing (NLP) methods, which are difficult to efficiently process complex multi-modal data (such as text, table, image). The bidding field has a unique and rigorous professional terminology system and legal and regulatory framework. Ordinary data recognition technology lacks in-depth learning and understanding of these professional knowledge, and is prone to misinterpretation and misjudgment when identifying and processing related content, affecting the judgment of the legality and compliance of the bidding documents. In addition, existing data recognition technology lacks field-adapted fine-tuning methods. Bidding field data is very different from general field data. Most current data recognition technologies are based on general field data training and optimization, and do not fully consider the special nature of the bidding field, making it difficult to adapt to data characteristics when applied to this field, with low information extraction accuracy.

[0004] Therefore, there is an urgent need to provide a large language model suitable for the bidding field, which has the ability to understand knowledge and process data. SUMMARY

[0005] To overcome the problems in the related art, the present disclosure provides a large language model construction method, system, device and medium for identifying bidding data to solve the technical problems in the related art.

[0006] One or more embodiments of the present specification provide a large language model construction method for identifying bidding data, comprising the steps of: Obtaining multi-modal bidding data, identifying the structured data of each modal data through information extraction technology, obtaining the text content and semantic labels of the corresponding type of text content of each data, and performing data format standardization processing on the processed multi-modal bidding data to obtain cross-format text data; Based on the cross-format text data, all legal terms are filtered and obtained, and a basic term dictionary is constructed according to the legal term version. The semantic association relationship between each legal term is determined according to the basic term dictionary combined with the restart random walk algorithm, and a legal term knowledge graph is constructed; Based on the basic terminology dictionary, the legal terminology level in the basic terminology dictionary is determined according to the constructed legal terminology level evaluation system, and the weight coefficient of the legal terminology is determined through a preset time decay factor; then the legal terminology score of each legal terminology is determined according to the legal terminology level, the weight coefficient of the legal terminology and a preset legal terminology credibility scoring function, and the basic terminology dictionary is obtained; and based on the basic terminology dictionary, a legal terminology field dictionary is constructed; The legal provision text data and the case precedent text data are acquired, and the legal provisions are cleaned of redundant information; then the legal provisions are standardized based on a Gini purity filter, and the processed legal provisions are obtained; the citation relationship of the case precedents is determined based on the case precedent text, and a directed acyclic legal provision network graph is constructed; Based on the preset provision level weight, the professional domain weight and the historical precedent rate of each legal provision, the effectiveness value of each processed legal provision is determined as the value of each node; the decay degree of the legal provision effectiveness with time or scene is quantified based on a constructed multi-factor effectiveness decay formula, the effectiveness propagation relationship between the legal provisions is determined, the edge weight is obtained, and a weighted provision network is obtained; The overlap degree between the legal provisions or between the case precedents in the case precedent text data is determined, the conflict score of the legal provisions is determined according to the effectiveness value of each legal provision and through a conflict detection function, and the legal provisions with conflicts are determined; according to the determined conflict judgment result, the keywords of the conflicting legal provisions are determined by using regular matching, and the final output of the weighted provision network display is displayed according to the citation relationship and a risk report on the conflict risk is generated; wherein the judgment priority constraint is that the judgment effectiveness of the higher law is higher than that of the lower law; Based on the weighted provision network, the bidirectional citation relationship between the case precedents and the processed legal provisions is determined in combination with the case precedent text data, the annotated case-provision association pair is generated, the bidirectional weight of the annotated case-provision association pair is calculated by using a bidirectional attention mechanism, the attention matrix is generated, and finally a bidirectional weighted association graph is generated; the legal terminology weight in the optimized basic dictionary is extracted by using the bidirectional weighted association graph, the incremental learning and fine-tuning training of the large language model are realized by using model fine-tuning, and a large language model suitable for bid and tender data recognition is obtained.

[0007] One or more embodiments of the present specification provide a large language model construction system for identifying bid and tender data, which comprises a terminology processing layer, a legal provision processing layer, a case decision processing layer and a model fine-tuning layer; The terminology processing layer comprises a document format processing module, a legal terminology knowledge graph construction module and a legal terminology field dictionary construction module; The document format processing module is configured to acquire multi-modal bidding data, identify structured data of each modality data through information extraction technology, obtain text content corresponding to each data and semantic labels corresponding to the type of the text content, and obtain cross-format text data. The legal term knowledge graph construction module is configured to filter all legal terms based on the cross-format text data, construct a basic term dictionary based on legal term versions, determine semantic association relationships between the legal terms according to the basic term dictionary and a restart random walk algorithm, and construct a legal term knowledge graph. The legal term field dictionary construction module is configured to determine the levels of the legal terms in the basic term dictionary according to a constructed legal term level evaluation system based on the basic term dictionary, determine weight coefficients of the legal terms through a preset time decay factor, determine scores of the legal terms according to the levels of the legal terms, the weight coefficients of the legal terms, and a preset legal term credibility scoring function, perform legal term retention filtering based on the scores of the legal terms, and obtain an optimized basic term dictionary; and construct a legal term field dictionary based on the optimized basic term dictionary. The legal clause processing layer includes a legal clause network graph construction module and a weighted clause network construction module. The legal clause network graph construction module is configured to acquire legal clause text data and case precedent text data, perform redundancy information cleaning on legal clause elements based on legal clauses, perform standardization processing on the legal clauses based on a Gini purity filter, determine citation relationships of precedents based on the case precedent text, and construct a directed acyclic legal clause network graph. The weighted clause network construction module is configured to use the constructed multi-factor effectiveness decay formula to quantify the decay degree of legal clause effectiveness over time or scenarios when conflicts exist between two legal clauses, determine effectiveness propagation relationships between the legal clauses, obtain edge weights, and thus obtain a weighted clause network. The case decision processing layer includes a legal clause conflict score determination module, a weighted clause network optimization module, and a weighted association graph generation module. The legal clause conflict score determination module is configured to determine the overlap degree between legal clauses or between cases and precedents in the case precedent text data, determine the conflict scores of the legal clauses according to the effectiveness values of the legal clauses and through a conflict detection function, and thus determine legal clauses with conflicts. The weighted clause network optimization module is configured to determine the conflict decision results of the legal clauses according to the conflict scores of the legal clauses determined by the legal clause conflict score determination module, determine keywords of the conflicting legal clauses through regular matching, constrain the final output and display of the weighted clause network showing the citation relationships based on decision priority, and generate a risk report on conflict risks; wherein the decision priority constraint is that the decision effectiveness of a higher-ranking law is higher than that of a lower-ranking law. The weighted association graph generation module is used to determine the bidirectional citation relationship between case precedent texts and processed legal clauses based on the weighted clause network and case precedent text data, and generate labeled case-clause association pairs. The labeled case-clause association pairs are then used to calculate bidirectional weights using a bidirectional attention mechanism to generate an attention matrix, and finally generate a bidirectional weighted association graph. The model fine-tuning layer includes a model parameter fine-tuning module, which is used to extract the weights of each legal term in the optimized basic dictionary using a bidirectional weighted association graph. The model fine-tuning is used to achieve incremental learning and fine-tuning training of the large language model, so as to obtain a large language model suitable for bidding data recognition.

[0008] This specification provides one or more embodiments of a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the large language model construction method based on the identification of bidding data as described above.

[0009] This specification provides one or more embodiments of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for constructing a large language model for recognizing bidding data as described above.

[0010] This disclosure provides a method, system, device, and medium for constructing a large language model to identify bidding data. Its advantages lie in solving the efficiency, accuracy, and adaptability problems in traditional bidding data identification by integrating multimodal data processing and efficient parameter fine-tuning techniques. Furthermore, through the analysis and mining of information from bidding documents and case precedents, it helps improve the regulatory capabilities of bidding and procurement, addresses the one-sidedness of "information silos," and makes information more accurate and comprehensive. The acquisition method is simplified, effectively solving practical problems in bidding and procurement. Moreover, by analyzing large amounts of bidding data and based on the effectiveness and dissemination of defined clauses, it is possible to gain insights into market supply and demand, industry trends, and competitive landscape. For example, quantifying the actual binding force of certain clauses on the market allows companies to flexibly deal with weakly binding clauses, while highly binding clauses will eliminate suppliers who cannot meet the requirements. Additionally, statistical analysis of keyword frequency in some effectiveness metrics can predict future demand-side policy trends. Attached Figure Description

[0011] In order to make one or more embodiments of the present specification or the technical solutions in the prior art clearer, the drawings needed to be used in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0012] Figure 1 A flowchart of a method for constructing a large language model for identifying bidding data is provided for one or more embodiments of the present specification. Figure 2 A schematic diagram of multi-modal data classification identification is provided for one or more embodiments of the present specification. Figure 3 A constructed legal term knowledge graph schematic diagram is provided for one or more embodiments of the present specification. Figure 4 A schematic diagram of different versions of the same term is provided for one or more embodiments of the present specification. Figure 5 A citation relationship diagram of a case is provided for one or more embodiments of the present specification. Figure 6 A constructed weighted clause network schematic diagram is provided for one or more embodiments of the present specification. Figure 7 A system block diagram of a large language model for identifying bidding data is provided for one or more embodiments of the present specification. Figure 8 A structural schematic diagram of a computer device is provided for one or more embodiments of the present specification. DETAILED DESCRIPTION

[0013] In order to make one or more embodiments of the present specification or the technical solutions in the prior art clearer, the drawings needed to be used in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0014] The present application will be described in detail below with reference to specific embodiments and the accompanying drawings.

[0015] Method embodiments According to an embodiment of the present application, a method for constructing a large language model for identifying bidding data is provided, as shown in Figure 1As shown, a large language model construction method flowchart for identifying bidding data is provided in this embodiment. According to the large language model construction method for identifying bidding data in the embodiment of the present application, the method comprises the following steps. In step S1, document format processing: obtaining multi-modal bidding data, identifying the structured data of each modal data through information extraction technology, obtaining the text content and semantic labels corresponding to the type of the text content of each data, performing data format standardization processing on the processed multi-modal bidding data, and obtaining cross-format text data.

[0016] In this embodiment, the bidding data can include bidding announcements, bidding documents, technical solutions, commercial quotations, expert review opinions, and scoring records, etc. The data formats can be PDF, Word, or scanned documents, etc. The text content corresponding types include text, table, image, etc. The information extraction technology can combine Optical Character Recognition (OCR), Bidirectional Encoder Representation from Transformers (BERT), LayoutLM model, etc. The LayoutLM model is suitable for processing text + layout information in PDF and scanned documents.

[0017] The first element list of the structured data is obtained by sequentially processing each image through the OCR technology. The first element list includes element text content, element position coordinates, semantic paragraph labels, and reading order data.

[0018] The text content and position box coordinates output by the OCR are input into the layout analysis model together to realize end-to-end document structure understanding, and the text content and semantic labels corresponding to the type of the text content are output synchronously, such as title, body, table, etc.

[0019] In an embodiment, step S1 specifically comprises the following steps: S11, based on the obtained multi-modal bidding data, analyzing the document types of each data, and selecting the corresponding recognition technology to extract document information; wherein the document types include text documents, scanned documents, and text + image mixed documents, for example, the text documents can include scanned files (jpg files) obtained by scanning paper files, and can also include PDF documents converted from word documents and images taken, etc.

[0020] In this embodiment, the unified processing entrance of the constructed multi-modal data can route the input document to the corresponding processing channel through the data document type automatic distribution mechanism. For reference Figure 2As shown, the schematic diagram of the multi-modal data classification recognition provided by the embodiment is shown, and specifically, the document information is extracted by selecting the corresponding recognition technology, including the following steps: S111, for the obtained multi-modal bidding data, judging the data type, if it is a text type document, directly extracting the original content, if it is a scanned document, enabling dynamic resolution OCR technology to identify, obtaining an element list; if the data is text type + image type, the text and image in the document are identified separately.

[0021] The result of the OCR technology to identify the scanned document is not limited to simple text content, but also includes rich metadata such as position information, confidence score, language information, formatted information, and table data, etc. These additional information enables the OCR technology to play a greater role in various application scenarios, such as automated document processing, data entry, image search, etc.; wherein the image includes a scanned image; After the text + image mixed type document is automatically split into text, table and image regions, the divide-and-conquer processing is performed, wherein the image part is identified by the OCR technology, because the position information can be determined, avoiding the information confusion caused by the typesetting error.

[0022] Further, through the unified processing entrance of the multi-modal data constructed, the input document can be routed to the corresponding processing channel through the data document type automatic distribution mechanism, which can be specifically: A unified processing entrance based on file type automatic distribution is constructed, and the intelligent routing of the processing path is realized through file extension detection.

[0023] The modular design is adopted to encapsulate the recognition methods of different file types as standardized methods (_process_image: image processing flow / _process_pdf: PDF processing flow / _process_word: word processing flow), and a consistent calling mode and data structure are provided to the external system, and an automatic type discrimination mechanism of PDF document is innovatively realized to intelligently distinguish text type PDF and scanned image.

[0024] In the embodiment, the multi-modal bidding data may include text type PDF and scanned image, and there is a resolution sensitivity problem in processing the text type PDF and scanned image data. The traditional method is difficult to balance between processing quality and performance. The traditional static resolution (only using high resolution or low resolution) setting has the following problems: Although high resolution can improve the recognition accuracy of OCR, it may cause the data volume to grow exponentially, which makes the processing time grow continuously, while low resolution indeed enhances the processing speed, but when the text itself is not very clear or there are some linear table structures, it cannot be recognized or recognized accurately, so a balance between the two needs to be found.

[0025] Therefore, before step S11, the following steps are further included: Step S10, based on the established resolution selection gradient rule, dynamically adjusting the processing resolution of OCR by pre-analyzing the page element characteristics of the scanned document, and using the corresponding resolution to recognize the corresponding text type scanned document or image type scanned document, and the specific resolution selection gradient rule is: page elements are text dominant type: 250-300 dpi , page elements are text + picture mixed type: 300-350 dpi , page elements are image / bill type: 350-400 dpi ; preferably,

[0026] The resolution selection gradient rule set by the embodiment can automatically adjust the recognition resolution of OCR according to the page elements of the document, so as to ensure the accuracy of the recognition result.

[0027] S12, using the fusion of OCR recognition and document element layout analysis to analyze multi-modal document data, and synchronously outputting the structured data of the structured semantic label of each image, wherein the structured semantic label includes element position coordinates and semantic paragraph labels.

[0028] The existing OCR system can only extract text content and cannot retain the layout structure and semantic information of the document, resulting in difficulty in subsequent information extraction. The present patent fuses OCR text recognition and document element layout analysis to create a "text-coordinate-reading order" joint input mechanism: element layout analysis is performed on the text content and position box coordinates output by OCR to realize end-to-end document structure understanding and obtain a corresponding element list. The element list can include data such as element text content, element position coordinates, semantic paragraph labels, and reading order. The text content and corresponding semantic labels (title / body / table, etc.) are synchronously output. The technical content is known to those skilled in the art, and will not be described in detail here.

[0029] S13, based on the position coordinate information of each element, restructuring the processing results of various types of documents, uniformly converting into cross-format structured data in a standardized legal analysis special data format, and outputting the cross-format structured data; The technical scheme realizes efficient processing of legal documents in the following ways: first, using intelligent recognition technology to automatically distinguish between different formats of documents such as PDF, Word, and image, to ensure that all types of documents can be correctly processed. Then, combining text recognition and layout analysis technology, not only the document content is extracted, but also the original format and position information is retained. Finally, the output data adopts a unified standard format and can be directly used for subsequent legal analysis.

[0030] This processing method significantly improves work efficiency and automates the document processing process that originally required manual operation. At the same time, the standardized output format makes the analysis results more reliable, providing a solid data foundation for legal risk assessment.

[0031] The embodiment also includes confidence checking of the recognition result information of S12 through the set double confidence checking mechanism. In the traditional recognition mode, there is a single confidence trap. The double confidence checking mechanism is set to comprehensively check the OCR recognition confidence and the element layout analysis result for quality evaluation, including confirmation of the confidence of each character in the OCR character recognition result and confirmation of the position of each element in the structured data. In a specific embodiment, the first confidence checking mechanism is OCR character-level confidence checking, which determines the accuracy of each recognized character based on a preset confidence threshold, avoids falling into a local optimal solution and ignoring the global solution, and the second confidence checking mechanism is to compare the position coordinates of each element determined by OCR recognition with the position information of the corresponding elements in the structured data obtained from the recognition result. For inconsistencies, manual correction can be performed.

[0032] The embodiment realizes efficient processing of legal documents through the above-mentioned manner. First, intelligent recognition technology is used to automatically distinguish between different formats of documents such as PDF, Word, and scanned copies, ensuring that all types of files can be correctly processed. Then, combined with text recognition and layout analysis technology, not only the document content is extracted, but also the original format and structure information is preserved. The system will automatically adjust the processing parameters according to the document quality to ensure the accuracy of the recognition results. The final output data adopts a unified standard format and can be directly used for subsequent legal analysis. This processing method significantly improves work efficiency and automates the document processing process that originally required manual operation. At the same time, the standardized output format makes the analysis results more reliable, providing a solid data foundation for legal risk assessment.

[0033] Step S2, legal terminology knowledge graph construction: based on cross-format text data filtering to obtain all legal terminologies, and based on legal terminology version to construct a basic terminology dictionary, according to the basic terminology dictionary combined with the restart random walk algorithm (RWR) to determine the semantic association relationship between legal terminologies, and construct a legal terminology knowledge graph, which can be specifically referred to as Figure 3 The legal terminology knowledge graph constructed by the embodiment is shown in the schematic diagram of the legal terminology knowledge graph provided by the embodiment.

[0034] The same legal requirements are expressed in the legal terms of this embodiment. There may be multiple source heterogeneous challenges of national standard terms, local variants, and coexistence of historical versions, and there are certain relationships between legal terms in different scenarios, such as “breach of contract liability” associated with “compensation amount” and “statute of limitations”. Therefore, this embodiment uses distributed technology and knowledge graph method to comprehensively collect and associate various legal terms and their variants to form a legal term network system, and to lay the foundation for subsequent steps to judge the timeliness and legal effectiveness of the dynamic changes of a legal term.

[0035] In the bidding neighborhood, referring to Figure 4 As shown, the same term may have multiple source heterogeneous challenges of national standard terms, local variants, and coexistence of historical versions. To solve this problem, the authority and reliability of a term in a specific context need to be quantitatively evaluated. First, determine the versions of legal terms, including national standard legal terms, local variant legal terms, and historical version legal terms. Then, use the random walk algorithm (RMR) to mine the implicit association between legal terms, and build a legal term knowledge graph. Therefore, step S2 includes the following steps: S21, legal term data preprocessing and classification: based on cross-format text data, traverse and detect whether each legal term in each cross-format text is in the basic term dictionary. If it exists, it is not processed. If it does not exist, determine the corresponding legal term version and store it in the basic term dictionary, thereby obtaining the basic term dictionary.

[0036] S22, semantic association analysis and knowledge graph construction: based on the basic term dictionary, use the restart random walk algorithm (RWR) to automatically identify the semantic association relationship between standard terms, local variants, and historical versions, standardize language expression, construct a term knowledge graph, and form a complete legal term network system.

[0037] In this embodiment, the relationships between each legal term in the basic term dictionary are not reflected, and each term is isolated from each other. Therefore, indirect associations between legal terms need to be mined (such as the big model does not know what the relationship between the bidding document and the prequalification is. Based on the bidding text legal terms, it is determined that “bidding document” is associated with “prequalification” through “bidder”, which is the core term association density, thereby realizing the term knowledge graph. Step S3, legal term field dictionary construction: based on the basic term dictionary, the legal term level in the basic term dictionary is determined according to the constructed legal term level evaluation system (three levels), and the weight coefficient of the legal term is determined through the preset time decay factor; then the score of each legal term is determined according to the legal term level, the weight coefficient of the legal term and the preset legal term credibility scoring function, and the legal term reservation screening is carried out according to the legal term score, so as to obtain the optimized basic term dictionary; based on the optimized basic term dictionary, the legal term field dictionary is constructed.

[0038] Among them, the legal term version includes national standard legal term, local variant legal term and historical version legal term; the legal term level evaluation system includes authority level, general level and time limit level, wherein the legal term with high frequency of national standard is the authority level, the industry generally recognized is the general level, and the historical version and the new version legal term is the time limit level.

[0039] In order to synchronize and integrate multiple data, the legal term h credibility score (TrustScore) function is innovatively put forward, which is more skilled in capturing extreme values:

[0040] Among them, The national term standardization weight formula is as follows: ; ; Among them, The preset version weight of legal term h is The tender document set is T, and T is the total frequency of all legal terms appearing in the document, wherein the version weight of legal term h is proportional to the promulgation time, the closer the promulgation time is to the current time, the higher the weight is, and vice versa, for example, if the promulgation time of a legal term is 2025, the weight is set to 1, if the promulgation time of a legal term is 2017, the weight decreases to 0.75, C(t) is the original number of times of legal term h appearing in the national standard; Max t∈T C t ​​) is the number of times of the legal term h with the highest frequency among all legal terms, for example, in the state law or the case, the term is not a definition class, such as the tender document explains that "the law stipulates that xxxx shall be subject to xxxx liability for breach of contract xxxx shall be subject to the obligation of the tenderer xxxx, otherwise the tenderer shall be subject to xxxx liability for breach of contract, and the tenderer shall also xxxx, in order to ensure that the tenderer xxxx", among them, the legal term "tenderer" appears 3 times, the legal term "breach of contract" appears 2 times, and the legal term "tenderer" appears the most times among all terms and the number of times is 3.

[0041] The dynamic factor a is calculated as follows: ; Among them, is the validation set accuracy of the national standard legal term, is the learning rate, is the validation set accuracy of the local variant legal term, which can be set artificially, for example, 1000 national standard legal terms are input into the model for classification, and the model finally identifies 800 national standard legal terms, so that is 0.8, is determined in the same way.

[0042] In order to ensure the form standardization of legal terms and reduce the influence of differences, the maximum normalization method is used here to eliminate the deviation of the relative importance of legal terms caused by different standard document lengths, and the frequency of the corresponding legal term in the corresponding document is converted into a relative value in the interval [0, 1]. Because the length of each document is different, for example, document A is 3 pages of content, and the frequency of legal term A in the entire document A is 35 times, while document B is 300 pages of content, and the frequency of legal term A in the entire document B is 35 times, but the importance of legal term A for the two documents is obviously different, so the difference is eliminated and normalized. Because directly using the integer 3 and the integer 35 into the national term standardization weight formula, the influence on the calculation result of the national term standardization weight will be great, and it will not play a role in convergence control at all, therefore, the frequency of the legal term in the corresponding document is converted and normalized to 0-1, which can ensure the stability of the national term standardization weight.

[0043] But when the statistical frequency of a legal term is 0, zero value overflow is easy to occur, and Laplace smoothing is used in this embodiment to avoid zero frequency problem, the formula is as follows: ; wherein DF(h) is the frequency of legal term h in the document, and T is the total frequency of all legal terms in the document; The decay factor is used to process historical versions of legal terms, mainly to address the timeliness and dynamic changes in legal effectiveness of legal terms, and the formula is as follows: Preferably, the embodiment further includes a confidence test for each legal term score, and legal terms exceeding the confidence threshold are retained, otherwise they are manually reviewed to determine whether they are invalid and corrected, and if corrected, the corrected legal terms are determined by a preset time decay factor.

[0044] In an embodiment, the confidence function is as follows: .

[0045] The embodiment realizes the systematic management of legal terms through intelligent means, uses distributed technology and knowledge graph method, comprehensively collects and associates various legal terms and their variants, and realizes the whole process processing from data collection to analysis and evaluation. In the evaluation mechanism, a three-level authoritative evaluation system is innovatively established, and through the TrustScore calculation model, professional problems such as term standardization and version update are effectively handled. The final formed field dictionary not only includes standard terms, but also systematically records the historical evolution process and use scenarios of the terms. The application of the scheme significantly improves the quality and efficiency of legal term management. Through the automatic processing process, the manual operation link is greatly reduced, and the term management work is more standardized and unified. The standardized term dictionary provides a reliable basis for legal document processing, helps to improve the professionalism and consistency of legal services, and has a positive significance for promoting the digital transformation of the legal industry.

[0046] Compared with the traditional TF-IDF, there are three dimensions of improvement: in the evaluation standard, a multi-level authoritative evaluation system is innovatively constructed, and the national standard adoption, cross-agency use breadth and version timeliness are included in the unified framework. This improvement effectively solves the problem of high weight of non-standard terms in the industry (such as "envelope" and other non-standard expressions); in the parameter optimization, trainable weight coefficients are used instead of traditional manual setting, and dynamic adjustment is realized through self-adaptive learning algorithm. This method shows good adaptability in different application scenarios such as government bidding and bidding; in the timeliness processing, a dynamic decay mechanism is innovatively introduced to automatically identify and reduce the weight coefficient of outdated terms (such as "base bid guarantee" and other expressions in the old version of the specification). Compared with the traditional static evaluation model, the new method can more accurately reflect the timeliness characteristics of term effectiveness.

[0047] Step S4, legal provision network graph construction: obtain legal provision text data and case precedent text data, clean up redundant information outside legal provision elements based on legal provisions; then standardize legal provisions based on Gini purity filter, obtain processed legal provisions; determine the citation relationship of precedents based on case precedent text, and construct a directed acyclic legal provision network graph.

[0048] In this embodiment, step S4 includes steps of: S41, obtain legal provision text data and case precedent text data, and verify the legality of legal provisions. If the legal provision is illegal, it is eliminated. Legal verification is used to determine whether the legal provision belongs to current law and historical law, superior law and subordinate law, or newly revised law. If it is an abandoned provision or a non-existent provision, it will be eliminated.

[0049] S42, clean up redundant information of legal provisions based on Gini coefficient purity filter, effectively remove templated content, and retain core legal provision elements; standardize legal provision text based on Gini purity filter.

[0050] In an embodiment, a part of the legal provisions has repeated templates, and the redundant information needs to be cleaned up. The Gini-based purity filter is used for preprocessing, as follows: ; In the formula, f k ( s ) is the number of times the kth entity appears in the sentence s, F( s ) is the total number of entities in the sentence s, and K is the entity type. The text filtering rule is: ; In this document, the threshold value θ is 0.4.

[0051] S43, determine the citation relationship of precedents based on case precedent text data. For details, refer to Figure 5 , which is a schematic diagram of the citation relationship of precedents provided in this embodiment. Specifically, the restart random walk algorithm (RWR), regular expression (Regex), or rule engine (such as Prolog, Drools) can be used.

[0052] S44, based on the citation relationship of precedents and the standardized legal provision text data, construct a provision network and efficiency modeling, and form an implicit association with the citation relationship to construct a legal provision directed acyclic graph (DAG) network.

[0053] Step S5, weight clause network construction: based on the preset clause level weight, professional domain weight and historical case rate of each legal clause, the effectiveness value of each processed legal clause is determined as the value of each node, the attenuation degree of the effectiveness of the legal clause with time or scene is quantified based on the constructed multi-factor effectiveness attenuation formula, the effectiveness propagation relationship between the legal clauses is determined, and the edge weight is obtained, the effectiveness propagation relationship between the legal clauses is quantified, and the weighted clause network is obtained.

[0054] In step S5 of the embodiment, the processed legal clauses are determined according to the corresponding level weight of the legal level Specifically, it is set as: ; The legal level from high to low is in turn constitution, law, administrative regulations and local regulations.

[0055] The legal level effectiveness value of the processed legal clauses is calculated as follows: ; Among them, is the historical case rate of the legal clause c, wherein Precedence(c) = the number of case precedents citing the legal clause c / the total number of related field cases, the frequency and breadth of the clause c cited in judicial practice, reflecting the actual application situation and judicial recognition degree of the clause, the initial value is artificially set; is the professional domain weight of the legal clause c, which indicates the application strength of the legal clause in a specific field, and is 1 if applicable and 0 if not applicable, which is used to represent the frequency or importance of the clause cited in a specific field; for example, there is a term called insider trading prohibition clause in securities law, which is important in the financial industry and is mentioned many times, so the professional domain weight of the insider trading prohibition clause is 1, but in the construction industry, the insider trading prohibition clause is not applicable, and the corresponding professional domain weight is 0.

[0056] Since the legal system follows the principle of "superior law over inferior law" (such as constitution > law > administrative regulations), all clause level weights The highest proportion, the stability of the case support, reflecting the practical recognition degree of the clause, a higher β value ensures that the legal clause needs to be tested by judicial practice, avoiding "paper law", therefore, in the embodiment 0.6, 0.3, 0.1 respectively.

[0057] In step S5, the attenuation degree of the effectiveness of the legal clause with time or scene is quantified based on the constructed multi-factor effectiveness attenuation formula, the effectiveness propagation relationship between the legal clauses is determined, and the weight value is determined, wherein the multi-factor effectiveness attenuation degree is calculated as follows:

[0058] is a preset basic force coefficient:

[0059] wherein, is a network node i of a legal clause i , the hierarchy of the clause hierarchy can refer to the clause hierarchy weight , the hierarchy from high to low is constitution → law → administrative regulation → local regulation, wherein, the root clause (constitution) is set to 0, is a network node i of a legal clause i and a network node j of a legal clause j , is the distance between the two, is the number of nodes in the weighted clause network interval, is a decay factor, this model strictly follows the legal interpretation principle of "explicit one excludes others", and the decay effect of indirect reference is strengthened through a natural exponential function, in order to accelerate the decay trend, a natural exponential is introduced. Reference Figure 6 is shown, which is a schematic diagram of a weighted clause network constructed in this embodiment.

[0060] This embodiment constructs an intelligent legal clause analysis system, first automatically cleans and standardizes the original legal clause text, innovatively determines the force propagation relationship, can accurately quantify the force relationship between legal clauses, and in the processing process, both the authority difference of legal hierarchy and the timeliness factor are considered, so that the analysis result is more in line with the demand of legal practice.

[0061] In the legal analysis in the field of bidding, the traditional method has obvious shortcomings. Although the keyword matching technology is simple and fast, it cannot identify implicit reference relationship and cannot reflect the legal force hierarchy, the maintenance cost is high and it is difficult to adapt to the update of regulations, the machine learning model such as Legal-BERT can capture semantic association, but it needs a large amount of labeled data, the decision-making process is not transparent, and when processing long text, the limitation is larger, the innovative force propagation relationship model (weighted clause network) solves these problems through a dynamic force decay mechanism. This model uses a multi-factor decay formula to consider legal hierarchy, reference distance and timeliness, and the advantage of this weighted clause network is its excellent accuracy and interpretability, through the construction of the weighted clause network, legal force visualization can be realized. Traceability, conflict clause impact range analysis and simulation and prediction of the impact of the modified regulations according to the node corresponding to the regulations.

[0062] Step S6, legal clause conflict score determination: determine the overlap degree between legal clauses or between cases in the case case text data, determine the conflict score of the legal clause according to the validity value of each legal clause, and determine the legal clause with conflict by the conflict detection function.

[0063] In this embodiment, taking the chain of references from the "Bidding and Tendering Law" to the "Implementation Clause" to the "Local Clause" as an example, the model accurately reflects the legal principle of "higher law than lower law" of legal clauses. When a legal clause in the tender document conflicts with the higher law, the model not only identifies the conflict, but also quantifies the conflict intensity. There will be conflicts between legal clauses, and there will be a probability of conflict between legal clauses and higher laws. For example, if some local rules shorten the legal tender deadline, the conflict score between the two legal clauses will be high, the risk will be higher, and the correct judgment rate will decrease. The conflict detection function is as follows:

[0064] Wherein, i And j are the validity values of legal clauses i and j , L is the legal clause text, Jaccard is the similarity between legal clauses i and j , len is the length of the legal clause text.

[0065] In this embodiment, the formula of the validity propagation algorithm is: ; In the formula, Force( v i ) and Force( v j ) are the legal validity values of nodes i and j respectively, P( v j ) is the set of predecessor nodes, and BaseForce( v j ) is a constant term.

[0066] In this embodiment, the validity value of the high-level node in the weighted clause network is propagated to the lower-level node by the formula, and the legal validity value is inherited from the cited legal clause. The hierarchical logic is roughly as follows: Constitution layer → Law layer → Administrative Regulations layer.

[0067] Step S7, Optimization of the Network of Rights-Bearing Clauses: Based on the conflict resolution results of Step S6, the keywords of the conflicting legal clauses are determined using regular expression matching (keywords in the clauses), and the citation relationship is displayed in the final output network of rights-bearing clauses based on the priority constraint of the judgment, and a risk report on the conflict risk is generated; wherein, the priority constraint of the judgment is that the judgment effect of the superior law is higher than that of the judgment effect of the subordinate law. For example, the effect of the local clause is lower than that of the implementing clause. When the two conflict, the implementing clause shall take precedence as the bidding standard.

[0068] In one embodiment, the risk report may include explanations, alerts, or reminders based on the identified content, as well as the type of legal clause, and provide a feedback report based on the identification results.

[0069] Step S8, Weighted Association Graph Generation: Based on the weighted clause network and combined with case precedent text data, determine the bidirectional citation relationship between the case precedent text and the processed legal clauses, and generate labeled case-clause association pairs. Calculate the bidirectional weights of the labeled case-clause association pairs using a bidirectional attention mechanism to generate an attention matrix, and finally generate a bidirectional weighted association graph.

[0070] In one embodiment, the bidirectional attention mechanism calculates the bidirectional weights as follows: ; in, Let n be the query matrix of case precedent texts, n be the length of the case precedent texts, and d be the attention head dimension. It is a key matrix of legal clauses, encoding the text and context of the legal clauses (m is the length of the clause), which is essentially vector processing; It is a value matrix of legal clauses, and Shared input but independent projection, where d is the attention head dimension. It is a scaling factor to prevent the gradient from vanishing due to excessively large numerator dot products. This is a reference matrix, calculated as follows:

[0071] The probabilistic conflict correction function is calculated as follows: ; in, Legal terms for time t i and j similarity, For legal terms i and jThe conflict score is determined by a conflict judgment function.

[0072] The embodiment takes the chain of references of the "Bidding and Tenders Law" to the "Implementation Regulations" to the "Local Regulations" as an example, and the model accurately reflects the legal principle that "superior law is superior to inferior law". When the bidding document conflicts with the superior law, the model not only identifies the conflict, but also quantifies the conflict intensity.

[0073] The preferred embodiment further provides a conflict judgment function to avoid invalid bidding caused by legal clause conflicts. The conflict judgment function is: a

[0074] wherein Force( v a ) and Force( v b ) are the legal force values of the conflicting clauses, and τ is a threshold value to avoid excessive sensitivity, for example, if two laws are slightly different, the score difference will not be large, and misjudgment will not occur. a and b

[0075] The embodiment effectively distinguishes explicit references and implicit associations between case documents and legal clauses through a bidirectional attention mechanism. The system uses a "force value-conflict score" double-layer analysis framework to accurately quantify the constraint relationship between clauses and intelligently identify potential conflicts. Taking the conflict between Article 5 of the "Bidding and Tenders Law" and the local regulations as an example, the model can automatically calculate the force decay process, and when the conflict score exceeds the threshold value, it immediately marks the suggestion "delete local clauses that conflict with superior laws" in the risk report. Compared with traditional methods, it has three advantages: strengthening the legal reference characteristics through prior matrix; dynamic conflict scoring for risk classification; and visual analysis to improve result interpretability.

[0076] Step S9, model parameter fine-tuning: using the bidirectional weighted association graph to extract the weights of each legal term in the optimized basic dictionary, using model fine-tuning to realize incremental learning and fine-tuning training of the large language model, and obtaining a large language model suitable for bidding and tendering data recognition.

[0077] The embodiment can use SD_LoRA technology for incremental learning by reducing the number of parameters and improving training efficiency.

[0078] The preferred embodiment of the large language model fine-tuning also considers the content of the wind direction report, so the large language model fine-tuning specifically includes: The weights of each legal term in the optimized basic dictionary are extracted by using a bidirectional weighted association graph, the legal term weights are adjusted by SD-LoRA technology, dynamic rank decomposition is performed, so as to realize dynamic scaling of the adaptation strength of the SD_LoRA module according to the input legal term weight (for example: high weight term triggers greater parameter adjustment), output the first parameter set, and generate legal clause type labels and corresponding bottleneck parameters d based on the risk report t The second parameter set is generated by using L-Adapter technology, wherein the legal clause type includes prohibition type, authorization type, obligation type and definition type, the first parameter set and the second parameter set are parameter fused, the model parameters are obtained and the model is fine-tuned, and a large language model suitable for bid data recognition is obtained.

[0079] The prohibition type legal clause can search for keywords such as 'prohibit' and 'not allowed' to determine, the authorization type legal clause can search for keywords such as 'have the right to' and 'can', the obligation type legal clause can search for keywords such as'should' and'must', and the definition type legal clause can search for keywords such as 'defined as' and'may be'.

[0080] In the embodiment, the legal term weight is adjusted by SD-LoRA, and dynamic rank decomposition is specifically: term weight mapping, mapping the credibility score of the term to the preset rank interval, wherein high credibility terms are assigned high ranks, and low credibility terms are assigned low ranks, and the rank distribution is corrected in real time according to the actual prediction accuracy of the term by the large language model, wherein the rank is reduced when the prediction error is higher than the threshold, and the rank is improved when the error is lower than the threshold, in addition, the mapping relationship is smoothly transitioned by an S-shaped function, that is, the corresponding formula r _t Because the rank itself is not derivable, smoothing is required here to ensure the trainability of backpropagation, and the corresponding derivative formula is

[0081] In this embodiment, the dual-channel adapter uses a hierarchical heterogeneous parameter activation strategy. One channel uses the SD-LoRA technology to generate a corresponding first parameter set based on the term weight, and the other channel generates a corresponding second parameter set based on the risk report. The large language model is fine-tuned based on the first parameter set and the second parameter set, so that the high-weight legal terms in the bidding documents can be automatically identified and the risks can be prompted. SD-LoRA and L-Adapter as adapters can insert lightweight modules into pre-trained language models (such as BERT, GPT) for efficient fine-tuning of specific tasks (such as parameter generation) without modifying the original model parameters. Adapt the general language model to the vertical field of risk clauses, dynamically generate a compliant parameter set (such as adjusting the penalty rate, period, etc.) according to the clause label and context. In addition, the weight value of each legal term is determined using the association graph, which ensures the domain relevance of the term weight. Then, according to the input legal term weight, the adaptation strength of the SD_LoRA module is dynamically scaled to assign higher ranks (retain more features) to high-weight terms and lower ranks (compress parameters) to low-weight terms, avoiding the model from paying too much attention to low-frequency terms, improving generalization, and dynamically adjusting the rank (i.e. complexity) of the model parameter matrix according to the term weight to balance effectiveness and efficiency.

[0082] In an embodiment, the SD-LoRA technology is an improvement based on existing LoRA technology. The term weight is dynamically mapped to the rank space of LoRA, breaking through the limitation of traditional fixed rank, and the dynamic rank is as follows: ; Among them, the hyperparameter setting is based on The theoretical basis is that when the minimum rank is maintained to avoid overfitting, The projection matrix is decomposed as

[0083] Among them, is the dynamically selected basis vector, Realize dynamic weight perception, because Floor is not derivable, gradient exploration is not smooth, and cannot guarantee trainability, so the gradient calculation uses approximate derivation Handle the non-differentiable problem of the Floor function: .

[0084] In an embodiment, the improved design of Adapter is L-Adapter, that is, the legal clauses are divided into 4 categories (prohibition / authorization / obligation / definition), and heterogeneous L_Adapter is designed to compress the clause type related dimensions, and different bottleneck dimensions are used for each type of clause and by adjusting the bottleneck dimension, the influence strength of different clause types on the model is controlled, more accurate parameter allocation is realized, and overfitting of simple clauses or underfitting of complex clauses is avoided, wherein ; wherein, is the original dimension of the input vector.

[0085] The parameter sharing mechanism of the L-Adapter is as follows: A low =P t ·W shared ; A high =W shared ·Q t ; wherein, A low is the dimension reduction matrix, A high is the dimension increasing matrix, P t is the clause type projection matrix, W shared is the sharing matrix, Q t is the type compression matrix.

[0086] Finally, the residual connection output = input x + 0.01 x dimension increasing result, x is the original input layer input vector.

[0087] The method provided by the embodiment for identifying the large language model of the bidding data solves the problems of efficiency, accuracy and adaptability in the traditional bidding data identification by fusing multi-modal data processing and efficient parameter fine-tuning technology, and through analysis and mining of information in the bidding documents and case documents, helps to improve the supervision ability of bidding and procurement, solves the one-sidedness of "island information", makes the information more accurate and comprehensive, and the acquisition method is simplified, effectively solves the practical problems in bidding and procurement, and can also analyze a large amount of bidding data, based on the determined effect of each clause, the effect of the clauses can be determined, the supply and demand relationship, industry trend and competition pattern of the market can be understood, for example, some clauses can quantify some actual constraints on the market, for clauses with weak constraints, enterprises can flexibly deal with them, and for clauses with high efficiency, some suppliers who cannot meet the requirements will be eliminated, and some key word frequencies are counted, so the future demand end policy trend can be predicted.

[0088] System embodiment According to the embodiment of the present application, a large language model construction system for identifying bidding data is provided, as shown in the accompanying drawings. Figure 7 The large language model construction system for identifying bidding data provided in the present embodiment is shown in the accompanying drawings. According to the large language model construction system for identifying bidding data of the present embodiment, it comprises a term processing layer, a legal clause processing layer, a case decision processing layer and a model fine-tuning layer. The term processing layer comprises a document format processing module 10, a legal term knowledge graph construction module 20 and a legal term field dictionary construction module 30. The document format processing module 10 is used to acquire multi-modal bidding data, identify the structured data of each modal data through information extraction technology, obtain the text content and semantic labels corresponding to the type of the text content of each data, perform data format standardization processing on the processed multi-modal bidding data, and obtain cross-format text data.

[0089] The legal term knowledge graph construction module 20 is used to filter and acquire all legal terms based on the cross-format text data, construct a basic term dictionary based on the legal term version, determine the semantic association relationship between each legal term according to the basic term dictionary combined with the restart random walk algorithm (RWR), and construct a legal term knowledge graph.

[0090] The legal term field dictionary construction module 30 determines the level of each legal term in the basic term dictionary based on the constructed legal term level evaluation system, and determines the weight coefficient of the legal term through a pre-set time decay factor; then determines the score of each legal term according to the legal term level, the weight coefficient of the legal term and a pre-set legal term credibility scoring function, and performs legal term retention screening according to the legal term score, thereby obtaining an optimized basic term dictionary; and constructs a legal term field dictionary based on the optimized basic term dictionary.

[0091] The legal clause processing layer comprises a legal clause network graph construction module 40 and a weighted clause network construction module 50. The legal clause network graph construction module 40 is used to acquire legal clause text data and case precedent text data, clean up redundant information outside legal clause elements based on legal clauses; then performs standardization processing on the legal clauses through a Gini purity filter, obtains processed legal clauses; determines the citation relationship of precedents based on case precedent text, and constructs a directed acyclic legal clause network graph.

[0092] The weighted clause network construction module 50 is configured to determine the validity value of each processed legal clause as the value of each node based on the clause level weight, professional domain weight and historical case rate preset for each legal clause, quantify the attenuation degree of the validity of the legal clause over time or scene based on the constructed multi-factor validity attenuation formula, determine the validity propagation relationship between the legal clauses, and thus obtain the edge weight, and quantify the validity propagation relationship between the legal clauses, and thus obtain the weighted clause network.

[0093] The case decision processing layer includes a legal clause conflict score determination module 60, a weighted clause network optimization module 70 and a weighted association graph generation module 80. The legal clause conflict score determination module 60 is configured to determine the overlap degree between legal clauses or between case precedents in the case precedent text data, determine the conflict score of the legal clauses according to the validity value of each legal clause and through a conflict detection function, and thus determine the legal clauses in conflict.

[0094] The weighted clause network optimization module 70 is configured to determine the conflict decision result determined by the legal clause conflict score determination module 60, determine the keywords of the conflicting legal clauses by regular matching, and based on the decision priority constraint, finally output the weighted clause network showing the reference relationship and generate a risk report on the conflict risk; wherein the decision priority constraint is that the decision validity of the superior law is higher than that of the subordinate law, such as the validity of the local clause is lower than that of the implementation clause, and when the two conflict, the implementation clause is preferred as the bidding standard.

[0095] The weighted association graph generation module 80 is configured to determine the bidirectional reference relationship between the case precedent text and the processed legal clauses based on the weighted clause network and in combination with the case precedent text data, generate the annotated case-clause association pair, calculate the bidirectional weight of the annotated case-clause association pair by using the bidirectional attention mechanism, generate the attention matrix, and finally generate the bidirectional weighted association graph.

[0096] The model fine-tuning layer includes a model parameter fine-tuning module 90 configured to extract the weight of each legal term in the optimized basic dictionary by using the bidirectional weighted association graph, realize the incremental learning and fine-tuning training of the large language model by using the model fine-tuning, and obtain the large language model suitable for the bidding data recognition.

[0097] In another embodiment, the model parameter fine-tuning module 90 is specifically configured to: The weights of each legal term in the optimized basic dictionary are extracted by using a bidirectional weighted association graph, the legal term weights are adjusted by using an SD-LoRA technology, dynamic rank decomposition is performed, so as to realize dynamic scaling of the adaptation strength of the LoRA module according to the input legal term weights, output a first parameter set, generate a legal clause type label and a corresponding bottleneck parameter based on a risk report, generate a second parameter set by using an L-Adapter technology, wherein the legal clause type includes a prohibition type, an authorization type, an obligation type and a definition type, the first parameter set and the second parameter set are subjected to parameter fusion, model parameters are obtained and model fine-tuning is performed, and a large language model suitable for bid data recognition is obtained.

[0098] The embodiment of the present application is a system embodiment corresponding to the above-mentioned method embodiment, and the specific operation of each module processing step can be understood with reference to the description of the method embodiment, which will not be repeated here.

[0099] As shown in Figure 8 The present application also provides a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the method for constructing a large language model for identifying bid data in the above-mentioned embodiment when executing the computer program.

[0100] The present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to implement the method for constructing a large language model for identifying bid data in the above-mentioned embodiment.

[0101] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of each method can be included. Any reference to memory, storage, database or other medium used in each embodiment provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0102] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the relevant part can be referred to the part of the method embodiment. The above-described device and system embodiments are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to the actual needs. Those skilled in the art can understand and implement without creative labor.

[0103] In addition, each functional module in each embodiment of the present disclosure can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional modules.

[0104] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and the contents not described in detail in the specification of the present application are the known technology of those skilled in the art.

Claims

1. A method for constructing a large language model to identify bidding data, characterized in that, Including the following steps: Acquire multimodal bidding data, identify the structured data of each modality through information extraction technology, obtain the corresponding text content and semantic tags of the text content type, and perform data format standardization processing on the processed multimodal bidding data to obtain cross-format text data; All legal terms are obtained by filtering cross-format text data, and a basic terminology dictionary is constructed according to the versions of legal terms. The semantic relationships between legal terms are determined by combining the basic terminology dictionary with a restarted random walk algorithm, and a legal terminology knowledge graph is constructed. Based on the basic terminology dictionary and the constructed legal terminology level evaluation system, the level of each legal term in the basic terminology dictionary is determined, and the weight coefficient of the legal term is determined by a preset time decay factor; then, the score of each legal term is determined according to the legal term level, the weight coefficient of the legal term, and the preset legal term credibility scoring function, thus obtaining the basic terminology dictionary; based on the basic terminology dictionary, a legal terminology domain dictionary is constructed. Obtain legal clause text data and case precedent text data, and perform redundant information cleaning on the legal clauses; Then, the legal clauses are standardized using the Gini purity filter to obtain the processed legal clauses; the citation relationships of the precedents are determined based on the case texts, and a directed acyclic network graph of legal clauses is constructed. Based on the pre-defined clause hierarchy weight, professional domain weight, and historical precedent rate of each legal clause, the validity value of each legal clause after processing is determined, which serves as the value of each node. Based on the constructed multi-factor effectiveness decay formula, the degree of decay of the effectiveness of legal provisions over time or in different scenarios is quantified, the effectiveness propagation relationship between legal provisions is determined, and thus the edge weights are obtained, resulting in a network of rights-bearing provisions. Determine the degree of overlap between legal clauses or between case precedents in the case precedent text data, determine the conflict score of the legal clauses based on the validity value of each legal clause and the conflict detection function, and identify the legal clauses that are in conflict. Based on the determined conflict judgment results, the keywords of the conflicting legal clauses are identified using regular expression matching, and the citation relationship is displayed in the final output network of rights-bearing clauses based on the judgment priority constraint, generating a risk report on the conflict risk; wherein, the judgment priority constraint is that the judgment of the superior law has higher effect than the judgment of the subordinate law. Based on the network of rights-bearing clauses and combined with case precedent text data, the bidirectional citation relationship between case precedent texts and processed legal clauses is determined, and labeled case-clause association pairs are generated. The bidirectional weights of the labeled case-clause association pairs are calculated using a bidirectional attention mechanism to generate an attention matrix, and finally a bidirectional weighted association graph is generated. The weights of each legal term in the optimized basic dictionary are extracted using the bidirectional weighted association graph, and incremental learning and fine-tuning of the large language model are achieved through model fine-tuning to obtain a large language model suitable for bidding data recognition.

2. The method for constructing a large language model to identify bidding data as described in claim 1, characterized in that, Also includes: When multimodal bidding data includes scanned documents, based on the established resolution selection gradient rule, the OCR processing resolution is dynamically adjusted by pre-analyzing the page element features of the scanned documents. The corresponding resolution is then used to recognize either text-based or image-based scanned documents. The resolution selection gradient rule is as follows: for text-dominant page elements: 250-300. dpi Page elements are a mix of text and images: 300-350 dpi Page elements are images / tickets: 350-400 dpi .

3. The method for constructing a large language model to identify bidding data as described in claim 1, characterized in that, The formula for scoring the credibility of the legal terminology is as follows: ; ; in, As a measure of national terminology standardization weight, The version weight of the legal term h is preset, where the weight is directly proportional to the promulgation time, and C(t) is the original number of occurrences of the legal term h in the national standard; Max t∈T ( C ( t (+1) represents the number of times the legal term h appears most frequently among all legal terms. The document is a set of tender documents, where T represents the total frequency of all legal terms appearing in the documents. The formula for calculating the dynamic factor α is: ; in, It is the accuracy rate of the verification set of national standard legal terminology. It's the learning rate. It is the accuracy of the validation set for local variant legal terms, based on human settings during testing; Attenuation factor As shown in the following formula: 。 4. The method for constructing a large language model to identify bidding data as described in claim 1, characterized in that, The validity value of each legal clause after the determination and processing is calculated as follows: ; in, The corresponding hierarchical weights are pre-defined for legal clause c at the legal level. Precedence(c) represents the historical precedent rate of legal clause c, where Precedence(c) = the number of cases citing legal clause c / the total number of cases in related fields. The frequency and breadth of legal clause c being cited in judicial practice are initially set by the individual. The domain weight of legal clause c represents the strength of the legal clause's applicability in a specific domain; it is 1 if applicable and 0 if not.

5. The method for constructing a large language model to identify bidding data as described in claim 1, characterized in that, The constructed multi-factor efficacy decay formula is as follows: The preset base efficacy coefficient: in, For network nodes i The hierarchy For network nodes i and network nodes j The distance between them is the number of nodes in the network with weighted clauses. This is the attenuation factor.

6. The method for constructing a large language model to identify bidding data as described in claim 1, characterized in that, The collision detection function is as follows: in, i and j Legal terms i and j The effectiveness value, L For the legal terms and conditions, Jaccard For legal terms i and legal terms j similarity, len This refers to the length of the legal clause text.

7. The method for constructing a large language model for identifying bidding data as described in claim 4, characterized in that, The labeled case-clause association pairs are used to calculate bidirectional weights using a bidirectional attention mechanism, specifically as follows: ; in, Let n be the query matrix of case precedent texts, n be the length of the case precedent texts, and d be the attention head dimension. It is a key matrix of legal clauses, encoding the text and context of the legal clauses, where m is the length of the legal clauses; It is a value matrix of legal clauses, and Shared input but independent projection. Scaling factor This is a reference matrix, calculated as follows: The probabilistic conflict correction function is calculated as follows: ; in, Legal terms for time t i and j similarity, For legal terms i and j Conflict score.

8. The method for constructing a large language model for identifying bidding data as described in claim 7, characterized in that, The process of incremental learning and fine-tuning of a large language model using model fine-tuning to obtain a large language model suitable for bidding data recognition specifically includes the following steps: The weights of legal terms in the optimized basic dictionary are extracted using a bidirectional weighted association graph. The weights of legal terms are then fine-tuned using the SD_LoRA model, and dynamic rank decomposition is performed to dynamically scale the adaptation strength of the SD_LoRA module according to the input legal term weights, and the first parameter set is output. Then, based on the risk report, legal clause type labels and corresponding bottleneck parameters are generated, and a second parameter set is generated using L-Adapter technology; The first and second parameter sets are fused to obtain model parameters, and the model is fine-tuned to obtain a large language model suitable for bidding data recognition.

9. A system for constructing a large language model to identify bidding data, characterized in that, It includes a terminology processing layer, a legal clause processing layer, a case decision processing layer, and a model fine-tuning layer; The terminology processing layer includes a document format processing module, a legal terminology knowledge graph construction module, and a legal terminology domain dictionary construction module; The document format processing module is used to acquire multimodal bidding data, identify the structured data of each modality through information extraction technology, obtain the corresponding text content of each data and the semantic tags of the corresponding type of text content, and obtain cross-format text data; The legal terminology knowledge graph construction module is used to extract all legal terms based on cross-format text data, build a basic terminology dictionary based on the versions of the legal terms, determine the semantic relationships between the legal terms based on the basic terminology dictionary and the restarted random walk algorithm, and construct the legal terminology knowledge graph. The legal terminology domain dictionary construction module, based on the basic terminology dictionary and the constructed legal terminology level evaluation system, determines the level of each legal term in the basic terminology dictionary and determines the weight coefficient of the legal term through a preset time decay factor; then, it determines the score of each legal term based on the legal term level, the weight coefficient of the legal term, and the preset legal term credibility scoring function; and then, it performs legal term retention and screening based on the legal term scores to obtain an optimized basic terminology dictionary; based on the optimized basic terminology dictionary, it constructs a legal terminology domain dictionary. The legal clause processing layer includes a legal clause network diagram construction module and a rights-bearing clause network construction module; The legal clause network diagram construction module is used to acquire legal clause text data and case precedent text data, and to clean up redundant information beyond the elements of legal clauses based on the legal clauses. Then, the legal clauses are standardized using the Gini purity filter to obtain the processed legal clauses; the citation relationships of the precedents are determined based on the case texts, and a directed acyclic network graph of legal clauses is constructed. The module for constructing the network of rights-bearing clauses is used to determine the value of each node when there is a conflict between the two. Based on the constructed multi-factor effectiveness decay formula, it quantifies the degree of decay of the effectiveness of legal clauses over time or in a scenario, determines the effectiveness propagation relationship between legal clauses, obtains the edge weights, and thus obtains the network of rights-bearing clauses. The case decision processing layer includes a module for determining the legal clause conflict score, a module for optimizing the network of weighted clauses, and a module for generating a weighted association graph; The Legal Clause Conflict Score Determination Module is used to determine the degree of overlap between legal clauses or between case precedents in the case precedent text data. Based on the validity value of each legal clause and the conflict detection function, the conflict score of the legal clause is determined, thereby identifying the legal clauses that are in conflict. The rights-bearing clause network optimization module is used to determine the conflict judgment results determined by the legal clause conflict score determination module, identify the keywords of conflicting legal clauses using regular expression matching, and finally output and display the rights-bearing clause network with citation relationships based on the judgment priority constraint, and generate a risk report on conflict risk; wherein, the judgment priority constraint is that the judgment of the superior law has higher effect than the judgment of the subordinate law. The weighted association graph generation module is used to determine the bidirectional citation relationship between case precedent texts and processed legal clauses based on the weighted clause network and case precedent text data, and generate labeled case-clause association pairs. The labeled case-clause association pairs are used to calculate bidirectional weights using a bidirectional attention mechanism to generate an attention matrix, and finally generate a bidirectional weighted association graph. The model fine-tuning layer includes a model parameter fine-tuning module, which is used to extract the weights of each legal term in the optimized basic dictionary using a bidirectional weighted association graph. The model fine-tuning is used to achieve incremental learning and fine-tuning training of the large language model, so as to obtain a large language model suitable for bidding data recognition.

10. The method for constructing a large language model for identifying bidding data as described in claim 9, characterized in that, The model parameter fine-tuning module is specifically configured for: The weights of legal terms in the optimized basic dictionary are extracted using a bidirectional weighted association graph. The weights of legal terms are then fine-tuned using the SD_LoRA model, and dynamic rank decomposition is performed to dynamically scale the adaptation strength of the SD_LoRA module according to the input legal term weights, outputting the first parameter set. Then, legal clause type labels and corresponding bottleneck parameters are generated based on the risk report, and the second parameter set is generated using L-Adapter technology. The first and second parameter sets are fused to obtain the model parameters, and the model is fine-tuned to obtain a large language model suitable for bidding data recognition.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for constructing a large language model based on the identification of bidding data as described in any one of claims 1 to 8.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for constructing a large language model for identifying bidding data as described in any one of claims 1 to 8.