Contract content query method and system based on artificial intelligence
By splitting the contract text into text blocks, building a structural diagram and using word vectors and association tables, the problem of insufficient semantic understanding of contract queries in the existing technology is solved, and efficient and accurate contract content query is achieved.
Patent Information
- Application Number
- CN202510643942.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing contract query system cannot accurately understand the semantic and logical relationship of the contract content, and the accuracy, relevance and efficiency of query results need to be improved.
Split the contract text into multiple text blocks, perform semantic analysis and extract keywords, build a contract information structure chart, use pre-trained models to obtain word vectors to calculate proximity, create an association table, store historical query information, and extract query results based on the same frequency keywords and association tables.
It improves the speed and accuracy of contract content query, can quickly locate key information, provide comprehensive query results, and improve user query efficiency.
Smart Images

Figure CN120470095A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a contract content query method and system based on artificial intelligence. Background Art
[0002] Contracts are important legal documents in business activities, and the search and management of their content is crucial to business operations. Traditional contract search methods rely primarily on manual retrieval, which is inefficient and prone to errors. With the development of artificial intelligence, technologies such as natural language processing and knowledge graphs have provided new solutions for contract content search.
[0003] A similar prior art Chinese patent application with publication number CN113505303A provides a method and device for contract information query, including: obtaining contract information to be queried input by a user; the contract information to be queried includes keywords; querying whether the keywords support searching for the contract information to be queried in the ES browser; if searching for the contract information to be queried in the ES browser is supported, searching for a contract identity identifier containing the contract information to be queried in the ES browser based on the contract information to be queried, and querying the contract information corresponding to the contract identity identifier in the database based on the contract identity identifier; if searching for the contract information to be queried in the ES browser is not supported, searching for the contract information to be queried in the database, and querying the contract information containing the contract information to be queried based on the contract information to be queried.
[0004] A similar prior art includes a Chinese patent application with publication number CN113033197A, which provides a construction contract law query method and device thereof, including: collecting construction contract laws and regulations, and digitizing the construction contract laws and regulations to establish a construction contract law database; performing text segmentation and stop word removal on the construction contract laws and regulations based on natural language processing technology, and then calculating feature words through a word frequency inverse text algorithm; performing synonym expansion query on feature words through a self-built construction contract law common terminology vocabulary and continuous bag-of-words model; calculating the similarity of contract laws and regulations based on a vector space model and a cosine function improvement method to obtain the corresponding legal provisions in the "Construction Contract Conditions"; and integrating the entire database and query system into a local server or smart device.
[0005] However, most existing contract query systems can only achieve simple text matching and cannot accurately understand the semantics and logical relationships of the contract content. The accuracy, relevance and efficiency of the query results need to be improved.
[0006] Therefore, the present invention provides a contract content query method and system based on artificial intelligence. Summary of the Invention
[0007] This application provides an artificial intelligence-based contract content query method and system for efficiently and accurately querying contract content.
[0008] In a first aspect, the present application provides a contract content query method based on artificial intelligence, the method comprising:
[0009] Split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, including the first keyword and the second keyword, and construct a contract information structure diagram based on the keywords;
[0010] Obtaining a word vector for each keyword based on the pre-trained first model, calculating the proximity between each two keywords in the text block based on the word vectors, improving the contract information structure diagram based on the proximity, and creating a first association table;
[0011] Store historical query information, obtain query keywords based on real-time query statements, and search for keywords that appear at the same time as the query keywords based on historical query information;
[0012] The first information and the second information are extracted from the contract information structure diagram based on the query keyword, the same-frequency keyword and the first association table, and the query result is displayed based on the first information and the second information.
[0013] In conjunction with the first aspect, in a first implementation of the first aspect of the present application, constructing a contract information structure diagram based on keywords includes:
[0014] The first keyword is used as the key point of the contract information structure diagram to determine whether there is a correlation between each two first keywords. If so, a first connecting line is added to the two key points corresponding to the two first keywords with the correlation, and the correlation between the two first keywords is marked.
[0015] In combination with the first aspect, in a second implementation of the first aspect of the present application, improving the contract information structure diagram based on proximity includes:
[0016] Combining first keywords whose proximity is greater than a preset first threshold to generate a first keyword group, selecting a keyword from the first keyword group as a target keyword, and replacing the first keyword in the contract information structure diagram that belongs to the same first keyword group as the target keyword with the target keyword;
[0017] The second keywords whose proximity is greater than the first threshold are combined to generate a second keyword group, a third keyword is generated based on the second keyword group, corresponding key points are added to the contract information structure diagram based on the third keyword, and a second connecting line is added between the key points corresponding to the third keyword and the keywords in the second keyword group.
[0018] In combination with the first aspect, in a third implementation of the first aspect of the present application, improving the contract information structure diagram based on proximity includes:
[0019] Obtaining the number of first connecting lines connected to each key point, identifying the first keyword corresponding to the key point whose number of lines is greater than a preset second threshold as a characteristic keyword, and recording all characteristic keywords and their positions in the contract information structure diagram;
[0020] A standard vocabulary is created in advance, and the standard vocabulary stores standard words with high query frequency. For each keyword in the contract information structure diagram, the proximity between the keyword and the standard word in the standard vocabulary is calculated, and the first standard word with a proximity greater than a first threshold is obtained. The first standard word with the largest proximity is obtained as the second standard word, and the corresponding keyword is replaced with the second standard word. The second standard word and the keyword are also associated and saved in the first association table.
[0021] In combination with the first aspect, in a fourth implementation of the first aspect of the present application, searching for co-frequency keywords that appear simultaneously with the query keyword based on historical query information includes:
[0022] All keywords in the historical query information are referred to as fourth keywords, and the fourth keywords and query keywords are combined in pairs to generate multiple keyword pairs;
[0023] Divide the historical query information into multiple different sub-information sets according to time periods, extract historical keyword groups from each sub-information set, and for each keyword pair, count the number of first occurrences, second occurrences, and third occurrences of each historical keyword group simultaneously.
[0024] Compare the second number and the third number, obtain the maximum value, add all the first numbers to obtain the first result value, divide the first result value by the maximum value to obtain the second result value, obtain the fourth keyword corresponding to the second result value greater than the second threshold, and use the fourth keyword as the same-frequency keyword of the corresponding query keyword.
[0025] In combination with the first aspect, in a fifth implementation of the first aspect of the present application, extracting the first information and the second information from the contract information structure diagram includes:
[0026] searching for a standard word matching the query keyword from the first association table, and obtaining a candidate keyword based on the first association table if a matching standard word is found;
[0027] Determine whether there is a feature keyword that matches the candidate keyword. If so, obtain the location of the feature keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram;
[0028] Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the first information.
[0029] In conjunction with the first aspect, in a sixth implementation of the first aspect of the present application, when no matching standard word is found, the method includes: searching the first association table for a standard word that matches the same-frequency keyword; and when a matching standard word is found, obtaining a candidate keyword based on the first association table;
[0030] Determine whether there is a characteristic keyword that matches the same-frequency keyword. If so, obtain the location of the characteristic keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram;
[0031] Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the second information.
[0032] In combination with the first aspect, in a seventh implementation of the first aspect of the present application, displaying query results based on the first information and the second information includes:
[0033] Duplicate information in the first information and the second information is obtained, the duplicate information is placed in the first position in the query result, and other information other than the duplicate information is sorted and displayed based on the proximity to the query keyword.
[0034] In a second aspect, the present application provides an artificial intelligence-based contract content query system, the system comprising:
[0035] a structure diagram generating unit, configured to split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, the keywords including a first keyword and a second keyword, and construct a contract information structure diagram based on the keywords;
[0036] a structure graph updating unit, configured to obtain a word vector for each keyword based on a pre-trained first model, calculate the proximity between each two keywords in a text block based on the word vector, improve the contract information structure graph based on the proximity, and create a first association table;
[0037] A co-frequency keyword search unit is used to store historical query information, obtain query keywords based on real-time query statements, and search for co-frequency keywords that appear at the same time as the query keywords based on historical query information;
[0038] A result query unit is used to extract the first information and the second information from the contract information structure diagram based on the query keyword, the same-frequency keyword and the first association table, and display the query result based on the first information and the second information.
[0039] Compared with the prior art, the beneficial effects of the present invention are at least as follows:
[0040] In the technical solution provided by the present application, the contract text is split into multiple text blocks, each text block is semantically parsed to extract keywords, and a contract information structure diagram is constructed based on the keywords. The contract content is logically classified and associated, so that the key information of the query can be quickly located during the query, which greatly reduces the search scope and improves the query speed; by improving the contract information structure diagram and creating a first association table, the structure of the contract structure diagram is optimized so that it can more accurately express the semantic relationship; more comprehensive content is queried for users through the same-frequency keywords, thereby improving the efficiency of user queries; the query results are displayed based on the first information and the second information, which can not only query the content directly related to the query keyword, but also find other related information related to it, and sort the queried first information and second information based on proximity, to help users obtain more comprehensive, accurate and intuitive query results. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0042] Figure 1 This is a schematic diagram of an embodiment of a contract content query method based on artificial intelligence in an embodiment of the present application;
[0043] Figure 2 It is a contract information structure diagram generated based on a text block of a lease contract in an embodiment of the present application;
[0044] Figure 3 This is the updated contract information structure diagram in the embodiment of the present application;
[0045] Figure 4 This is a schematic diagram of an embodiment of an artificial intelligence-based contract content query system in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The embodiments of the present application provide a contract content query method and system based on artificial intelligence. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0047] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the contract content query method based on artificial intelligence includes:
[0048] Step S1: Split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, the keywords include a first keyword and a second keyword, and construct a contract information structure diagram based on the keywords.
[0049] Specifically, in order to improve the intelligence level of contract management and provide users with efficient and accurate query services, the contract text is first split into multiple text blocks. For example, a paragraph in the contract text can be used as a text block. Then, each text block is semantically parsed to extract keywords. For example, natural language processing technology can be used to extract keywords from the text block. Keywords include first keywords and second keywords. The first keyword refers to nouns, including the two parties to the contract, obligations, rights and clause names, etc. The second keyword refers to verbs, including payment, provision of certain services, etc., which can express the relationship between the two first keywords. Then, a contract information structure diagram is constructed based on the keywords. The specific generation method will be explained in detail later. The generated contract information structure diagram is stored in a graph database. The contract information structure diagram structures the contract content in the form of key points and lines, and logically classifies and associates the contract content. When querying, it can quickly locate the key information of the query, greatly reducing the search scope and improving the query speed.
[0050] Step S2: Obtain a word vector for each keyword based on the pre-trained first model, calculate the proximity between every two keywords in the text block based on the word vector, improve the contract information structure diagram based on the proximity, and create a first association table.
[0051] Specifically, in order to optimize the structure of the contract structure diagram so that it can express semantic relationships more accurately, we first obtain the word vector of each keyword based on a pre-trained first model, such as the Word2Vec model, and then use a similarity algorithm based on the word vector, such as cosine similarity, to calculate the proximity between each two keywords in the text block. Then, we optimize the contract information structure diagram based on the proximity. The specific optimization method will be explained in detail later. While optimizing the contract information structure diagram, we also create a first association table. The first association table stores the keywords in the contract content and the corresponding standard words, and also stores the position of the keywords in the contract information structure diagram. The first association table can assist in subsequent queries on the contract content.
[0052] Step S3: store historical query information, obtain query keywords based on real-time query statements, and search for co-frequency keywords that appear simultaneously with the query keywords based on the historical query information.
[0053] Specifically, in order to enhance the convenience and accuracy of user information query, the query statements entered by the user and the query time are associated and saved to generate historical query information. When the user subsequently enters a real-time query statement, the query keywords are extracted based on natural language processing technology, and then the same-frequency keywords that appear at the same time as the query keywords are obtained based on the historical query information. Same-frequency keywords refer to keywords that are frequently entered by users at the same time as the query keywords within a preset time period. For example, when a user inquires about breach of contract liability, he or she will also enter the amount of compensation, breach of contract notice period, etc. within a short period of time before and after. Then these keywords are the same-frequency keywords of the query keywords. The specific method of obtaining the same-frequency keywords will be explained in detail later. After obtaining the same-frequency keywords, more comprehensive content can be queried for users based on the query keywords and the same-frequency keywords, thereby improving the efficiency of user queries.
[0054] Step S4: extract the first information and the second information from the contract information structure diagram based on the query keyword, the same-frequency keyword, and the first association table, and display the query results based on the first information and the second information.
[0055] Specifically, the first information refers to the content information directly or indirectly related to the query keyword found from the contract information structure diagram based on the query keyword, and the second information refers to the content information directly or indirectly related to the query keyword found from the contract information structure diagram based on the same frequency keyword. Subsequently, the first information and the second information are displayed based on the proximity to the query keyword, so as to display comprehensive query results to users and facilitate users to intuitively view the query results.
[0056] In a specific embodiment, the contract information structure diagram is constructed based on keywords, and the following steps are further performed:
[0057] The first keyword is used as the key point of the contract information structure diagram to determine whether there is a correlation between each two first keywords. If so, a first connecting line is added to the two key points corresponding to the two first keywords with the correlation, and the correlation between the two first keywords is marked.
[0058] Specifically, in order to improve the efficiency of contract content query, a contract information structure diagram is constructed. First, the first keyword is used as the key point in the contract information structure diagram. Then, based on the correlation between the first keywords, a first connecting line is added between the key points corresponding to the two first keywords with the correlation. The correlation between the two key points (corresponding to the two first keywords) is also marked on the first connecting line, such as Figure 2 The figure shows a contract information structure diagram generated based on a text block of a lease contract. The generated contract information structure diagram is converted into graph data and stored in a graph database. The structured form of the graph data facilitates subsequent query and update.
[0059] In a specific embodiment, improving the contract information structure diagram based on proximity includes the following steps:
[0060] Combining first keywords whose proximity is greater than a preset first threshold to generate a first keyword group, selecting a keyword from the first keyword group as a target keyword, and replacing the first keyword in the contract information structure diagram that belongs to the same first keyword group as the target keyword with the target keyword;
[0061] The second keywords whose proximity is greater than the first threshold are combined to generate a second keyword group, a third keyword is generated based on the second keyword group, corresponding key points are added to the contract information structure diagram based on the third keyword, and a second connecting line is added between the key points corresponding to the third keyword and the keywords in the second keyword group.
[0062] Specifically, in order to reduce redundancy in the contract information structure diagram and improve the simplicity of the contract information structure diagram, the first keywords whose proximity is greater than the preset first threshold are combined to generate a group of first keyword groups. The first keyword group contains multiple semantically similar first keywords. A first keyword is selected from the first keyword group as the target keyword. For example, if a group of first keywords contains houses and real estate, one of the real estates can be selected as the target keyword, and all the houses in the contract information structure diagram are replaced with real estate.
[0063] In order to enrich the content of the contract information structure diagram and enable it to more accurately express the semantic relationship between words, the second keywords with a proximity greater than the first threshold are combined to generate a second keyword group. For example, one of the second keyword groups is payment and payment. A third keyword is generated based on the second keyword group. The third keyword can summarize the summary word of the second keyword group. For example, the generated third keyword is payment. The key point corresponding to the third keyword is added to the contract information structure diagram, and a second connecting line is added between the keyword in the second keyword group and the third keyword. The second connecting line represents that there is a summary relationship between the third keyword and the second keyword.
[0064] The above method can make the contract information structure diagram more comprehensive in expressing the information in the contract, avoid missing important semantic relationships, and quickly find relevant key points through the third keyword in subsequent queries.
[0065] In a specific embodiment, the contract information structure diagram is improved based on proximity, and the following steps are further performed:
[0066] Obtaining the number of first connecting lines connected to each key point, identifying the first keyword corresponding to the key point whose number of lines is greater than a preset second threshold as a characteristic keyword, and recording all characteristic keywords and their positions in the contract information structure diagram;
[0067] A standard vocabulary is created in advance, and the standard vocabulary stores standard words with high query frequency. For each keyword in the contract information structure diagram, the proximity between the keyword and the standard word in the standard vocabulary is calculated, and the first standard word with a proximity greater than a first threshold is obtained. The first standard word with the largest proximity is obtained as the second standard word, and the corresponding keyword is replaced with the second standard word. The second standard word and the keyword are also associated and saved in the first association table.
[0068] Specifically, in order to further improve the contract information structure diagram, the number of first connecting lines connected to each key point in the contract information structure diagram is obtained. The more lines there are, the more other first keywords have a correlation with the first keyword corresponding to this key point, which means that this first keyword is very important. Therefore, the first keyword corresponding to the key point with a number of lines greater than the second threshold is identified as a feature keyword. The feature keyword refers to important information in the contract. All feature keywords and the position of the feature keywords in the contract information structure diagram are recorded.
[0069] In order to improve the accuracy and readability of the contract information structure diagram, a standard word library is created in advance. The standard word library stores standard words with high query frequency. For each keyword in the contract information structure diagram, the proximity between the standard words in the word library is calculated and annotated. For each keyword, a first standard word whose proximity to the keyword is greater than a first threshold is obtained. The proximity is greater than the first threshold, indicating that the first standard word and the keyword are close enough to be synonymous. Then, the first standard word with the greatest proximity to the keyword is selected from the first standard words as the second standard word, and the corresponding keyword is replaced with the second standard word, such as Figure 3 As shown, it is the updated contract information structure diagram, the second annotation word refers to the standard word with the greatest proximity to the keyword, and finally the second standard word and the keyword are associated and saved in the first association table.
[0070] Since the standard words selected are the ones that are closest to the keywords and the standard words are in line with the query habits, the above method can retain the original information of the contract to the greatest extent and improve the query efficiency of the contract content by replacing the standard words with keywords.
[0071] In a specific embodiment, searching for keywords that appear simultaneously with the query keyword based on historical query information specifically includes the following steps:
[0072] All keywords in the historical query information are referred to as fourth keywords, and the fourth keywords and query keywords are combined in pairs to generate multiple keyword pairs;
[0073] Divide the historical query information into multiple different sub-information sets according to time periods, extract historical keyword groups from each sub-information set, and for each keyword pair, count the number of first occurrences, second occurrences, and third occurrences of each historical keyword group simultaneously.
[0074] Compare the second number and the third number, obtain the maximum value, add all the first numbers to obtain the first result value, divide the first result value by the maximum value to obtain the second result value, obtain the fourth keyword corresponding to the second result value greater than the second threshold, and use the fourth keyword as the same-frequency keyword of the corresponding query keyword.
[0075] Specifically, in order to improve the comprehensiveness of the query, after obtaining the query keyword, keywords that frequently appear with the current query keyword are obtained from the historical query information. These keywords are called co-frequency keywords. First, all keywords in the historical query information are extracted and these keywords are called fourth keywords. The fourth keywords and the query keywords are combined in pairs to generate multiple keyword pairs. Then, the historical query information is divided into multiple different sub-information sets according to time periods. For example, the historical query information input within every half hour or ten minutes is used as a sub-information set. The historical keyword group in each sub-information set is obtained. For the two keywords in each keyword pair, the first number of their simultaneous appearances in the historical keyword group is counted. The second number refers to the number of times the fourth keyword in the keyword pair appears in the historical keyword group, and the third number refers to the number of times the query keyword in the keyword pair appears in the historical keyword group. Then, the second number and the third number are compared to obtain the maximum value. A first number refers to the number of times the keyword pair appears simultaneously in a historical keyword group. All first numbers are added to obtain the total number of times the keyword pair appears simultaneously in all historical keyword groups. The first result value is divided by the maximum value to obtain a second result value. The fourth keyword in the keyword pair corresponding to the second result value greater than the second threshold is obtained, and the fourth keyword is used as the co-frequency keyword of the corresponding query keyword.
[0076] In a specific embodiment, extracting the first information and the second information from the contract information structure diagram specifically includes the following steps:
[0077] searching for a standard word matching the query keyword from the first association table, and obtaining a candidate keyword based on the first association table if a matching standard word is found;
[0078] Determine whether there is a feature keyword that matches the candidate keyword. If so, obtain the location of the feature keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram;
[0079] Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the first information.
[0080] Specifically, in order to quickly query the contract content that matches the query keyword, first search for the standard word that matches the query keyword from the first association table. Since the standard word is a word that conforms to the user's query habits, and the first association table stores all standard words related to the contract content, the corresponding standard word is first searched from the first association table. When the matching standard word is found, the candidate keyword is obtained based on the first association table. The candidate keyword refers to the keyword extracted from the contract text corresponding to the standard word. Since the feature keyword is an important keyword that is marked, it is first determined whether there is a feature keyword that matches the candidate keyword. In some cases, the feature keyword is obtained in the contract information structure. The position and corresponding key points in the composition, search for associated associated key points based on the connecting lines of the key points, mark the key points, associated key points and corresponding connecting lines found as the first information, if there is no such information, it means that the query keyword and the feature keyword do not match, then directly search for the key points that match the candidate keywords from the contract information structure diagram, and then search for associated associated key points based on the connecting lines of the key points, mark the key points, associated key points and corresponding connecting lines found as the first information, each key point and connecting line has a corresponding keyword, these keywords are the contract content that matches or is related to the query keyword, so this information is used as the first information.
[0081] It should be noted that, when searching for standard words that match the query keywords from the first association table, the search can be performed by comparing the proximity between the standard words and the query keywords in the first association table. For example, the proximity between the standard words and the query keywords in the first association table is calculated. If there is a standard word whose proximity to the query keywords is greater than the first threshold, it means that there is a marked word that matches the query keywords in the first association table. Then, the standard word with the largest proximity is obtained from the standard words that meet the conditions as the final standard word that matches the query keywords. If the proximity between all standard words and the query keywords is less than the first threshold, it means that there is no standard word that matches the query keywords.
[0082] In a specific embodiment, when no matching standard word is found, the following steps are specifically included:
[0083] Searching for a standard word that matches the same-frequency keyword from the first association table, and obtaining a candidate keyword based on the first association table if a matching standard word is found;
[0084] Determine whether there is a characteristic keyword that matches the same-frequency keyword. If so, obtain the location of the characteristic keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram;
[0085] Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the second information.
[0086] Specifically, if no standard word matching the query keyword is found, it means that there may be no content matching the query keyword in the contract content. However, in order to provide users with comprehensive query content, the same-frequency keyword is a keyword that appears frequently at the same time as the query keyword. It is speculated that the user may look for content related to the same-frequency keyword next. Therefore, based on the above steps, the second information matching the same-frequency keyword is searched from the contract information structure diagram.
[0087] In a specific embodiment, displaying query results based on the first information and the second information specifically includes the following steps:
[0088] Duplicate information in the first information and the second information is obtained, the duplicate information is placed in the first position in the query result, and other information other than the duplicate information is sorted and displayed based on the proximity to the query keyword.
[0089] Specifically, in order to enable users to intuitively obtain query results, duplicate information in the first information and the second information is obtained. Since the first information is the content directly matched by the query keyword, and the second information is the content that the user may query next, if there is duplicate information in the first information and the second information, it is inferred that the duplicate information may be the content that the user most wants to query. Therefore, the duplicate information is placed in the first position in the query results, that is, the front position, so that the user can see it intuitively. Then, other information except the duplicate information is sorted and displayed based on the proximity to the query keyword. Through the above steps, not only the content directly related to the query keyword can be queried, but also other related information related to it can be found, helping users to query more comprehensive and accurate results.
[0090] The above describes the contract content query method based on artificial intelligence in the embodiment of the present application. The following describes the contract content query system based on artificial intelligence in the embodiment of the present application. Figure 4 In the embodiments of the present application, an embodiment of the contract content query system based on artificial intelligence includes:
[0091] a structure diagram generating unit, configured to split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, the keywords including a first keyword and a second keyword, and construct a contract information structure diagram based on the keywords;
[0092] a structure graph updating unit, configured to obtain a word vector for each keyword based on a pre-trained first model, calculate the proximity between each two keywords in a text block based on the word vector, improve the contract information structure graph based on the proximity, and create a first association table;
[0093] A co-frequency keyword search unit is used to store historical query information, obtain query keywords based on real-time query statements, and search for co-frequency keywords that appear at the same time as the query keywords based on historical query information;
[0094] A result query unit is used to extract the first information and the second information from the contract information structure diagram based on the query keyword, the same-frequency keyword and the first association table, and display the query result based on the first information and the second information.
[0095] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0097] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. The contract content query method based on artificial intelligence is characterized by: The method comprises: Split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, including the first keyword and the second keyword, and construct a contract information structure diagram based on the keywords; Obtaining a word vector for each keyword based on the pre-trained first model, calculating the proximity between each two keywords in the text block based on the word vectors, improving the contract information structure diagram based on the proximity, and creating a first association table; Store historical query information, obtain query keywords based on real-time query statements, and search for keywords that appear at the same time as the query keywords based on historical query information; The first information and the second information are extracted from the contract information structure diagram based on the query keyword, the same-frequency keyword and the first association table, and the query result is displayed based on the first information and the second information.
2. The method according to claim 1, characterized in that Build a contract information structure diagram based on keywords, including: The first keyword is used as the key point of the contract information structure diagram to determine whether there is a correlation between each two first keywords. If so, a first connecting line is added to the two key points corresponding to the two first keywords with the correlation, and the correlation between the two first keywords is marked.
3. The method according to claim 1, characterized in that Improve the contract information structure diagram based on proximity, including: Combining first keywords whose proximity is greater than a preset first threshold to generate a first keyword group, selecting a keyword from the first keyword group as a target keyword, and replacing the first keyword in the contract information structure diagram that belongs to the same first keyword group as the target keyword with the target keyword; The second keywords whose proximity is greater than the first threshold are combined to generate a second keyword group, a third keyword is generated based on the second keyword group, corresponding key points are added to the contract information structure diagram based on the third keyword, and a second connecting line is added between the key points corresponding to the third keyword and the keywords in the second keyword group.
4. The method according to claim 3, characterized in that Improved contract information structure diagram based on proximity, also including: Obtaining the number of first connecting lines connected to each key point, identifying the first keyword corresponding to the key point whose number of lines is greater than a preset second threshold as a characteristic keyword, and recording all characteristic keywords and their positions in the contract information structure diagram; A standard vocabulary is created in advance, and the standard vocabulary stores standard words with high query frequency. For each keyword in the contract information structure diagram, the proximity between the keyword and the standard word in the standard vocabulary is calculated, and the first standard word with a proximity greater than a first threshold is obtained. The first standard word with the largest proximity is obtained as the second standard word, and the corresponding keyword is replaced with the second standard word. The second standard word and the keyword are also associated and saved in the first association table.
5. The method according to claim 1, wherein Search for keywords that appear simultaneously with the query keyword based on historical query information, including: All keywords in the historical query information are referred to as fourth keywords, and the fourth keywords and query keywords are combined in pairs to generate multiple keyword pairs; Divide the historical query information into multiple different sub-information sets according to time periods, extract historical keyword groups from each sub-information set, and for each keyword pair, count the number of first occurrences, second occurrences, and third occurrences of each historical keyword group simultaneously. Compare the second number and the third number, obtain the maximum value, add all the first numbers to obtain the first result value, divide the first result value by the maximum value to obtain the second result value, obtain the fourth keyword corresponding to the second result value greater than the second threshold, and use the fourth keyword as the same-frequency keyword of the corresponding query keyword.
6. The method according to claim 1, characterized in that Extracting the first information and the second information from the contract information structure diagram includes: searching for a standard word matching the query keyword from the first association table, and obtaining a candidate keyword based on the first association table if a matching standard word is found; Determine whether there is a feature keyword that matches the candidate keyword. If so, obtain the location of the feature keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram; Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the first information.
7. The method according to claim 6, characterized in that If no matching standard word is found, execute: Searching for a standard word that matches the same-frequency keyword from the first association table, and obtaining a candidate keyword based on the first association table if a matching standard word is found; Determine whether there is a characteristic keyword that matches the same-frequency keyword. If so, obtain the location of the characteristic keyword in the contract information structure diagram and the corresponding key point. If not, find the key point that matches the candidate keyword in the contract information structure diagram; Based on the connection lines of the key points, the associated key points are searched for, and the found key points, the associated key points and the corresponding connection lines are used as the second information.
8. The method according to claim 1, characterized in that Displaying query results based on the first information and the second information includes: Duplicate information in the first and second information is obtained, the duplicate information is placed in the first position in the query result, and other information other than the duplicate information is sorted and displayed based on the proximity to the query keyword.
9. An artificial intelligence-based contract content query system, used to implement the artificial intelligence-based contract content query method according to any one of claims 1 to 8, characterized in that: The system comprises: a structure diagram generating unit, configured to split the contract text into multiple text blocks, perform semantic analysis on each text block to extract keywords, the keywords including a first keyword and a second keyword, and construct a contract information structure diagram based on the keywords; a structure graph updating unit, configured to obtain a word vector for each keyword based on a pre-trained first model, calculate the proximity between each two keywords in a text block based on the word vector, improve the contract information structure graph based on the proximity, and create a first association table; A co-frequency keyword search unit is used to store historical query information, obtain query keywords based on real-time query statements, and search for co-frequency keywords that appear at the same time as the query keywords based on historical query information; A result query unit is used to extract the first information and the second information from the contract information structure diagram based on the query keyword, the same-frequency keyword and the first association table, and display the query result based on the first information and the second information.
Citation Information
Patent Citations
Building construction contract and regulation query method and device
CN113033197A
Contract information query method and device
CN113505303A
Object query method and device based on keyword extraction, medium and equipment
CN112818091A
Keyword extraction method and device, computer equipment and storage medium
CN113656429A
Contract text intelligent analysis method based on deep data mining
CN114328822A