Efficient legal data searching method based on artificial intelligence

By analyzing the full permutations and combinations of target keywords and their reference keywords, and their relevance, legal data retrieval results are generated, solving the problem of incomplete retrieval in existing technologies and achieving more accurate and comprehensive legal document retrieval.

CN121501977AInactive Publication Date: 2026-02-10SHANDONG UNIV OF POLITICAL SCI & LAW
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511544879.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing keyword-based legal data search methods may result in incomplete search results due to differences in users' writing habits, affecting users' judgment of relevant laws.

Method used

By acquiring target keywords and their reference keywords input by target users, performing full permutations and combinations, analyzing the differences in the content of legal documents in different permutations, and combining the relevance and frequency of keywords, search results are generated to improve the comprehensiveness and accuracy of the search.

Benefits of technology

It effectively reduces the problem of missed detections, improves the comprehensiveness and accuracy of legal document retrieval, and ensures that the search results are not only based on the frequency of occurrence of the target keywords, but also take into account the related reference keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501977A_ABST
    Figure CN121501977A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a legal data efficient search method based on artificial intelligence, which comprises the following steps of: acquiring a target keyword input by a target user, marking keywords except the target keyword in all legal documents to be retrieved as reference keywords, and storing the reference keywords into a database; obtaining new target keywords and new reference keywords corresponding to the new target keywords according to the occurrence conditions of the reference keywords in all the to-be-retrieved legal documents which contain the target keywords and are in different permutation and combination forms; obtaining the correlation degree between each new target keyword and each new reference keyword corresponding to the new target keyword, and obtaining the correlation degree between each new target keyword and each new reference keyword corresponding to the new target keyword according to the occurrence frequency of each new target keyword and each new reference keyword in each legal document to be retrieved and the correlation degree between each new target keyword and each new reference keyword corresponding to the new target keyword; and the recommendation degree of each legal document to be retrieved is obtained, so that the retrieval result is generated for the target user, and the comprehensiveness of searching the legal data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an efficient method for searching legal data based on artificial intelligence. Background Technology

[0002] With the continuous development of society and the increasing societal societal sophistication of the legal system, new regulations and precedents are constantly emerging, leading to a surge in legal documents. This makes it increasingly difficult for individuals to fully grasp all legal resources. To improve the efficiency of legal document retrieval, enhance the quality of legal services, and accelerate decision-making, technologies utilizing artificial intelligence for efficient legal data searching are gaining increasing attention from various legal institutions and departments. Among these, keyword-based legal data search methods are the most basic and common type of retrieval technology.

[0003] Existing keyword-matching search methods primarily involve analyzing the relevance between legal documents and keywords by statistically analyzing the frequency of keywords in legal documents and considering factors such as document length. These relevance factors are then used to rank legal documents for search. However, due to differences in writing habits among users, the content expressed by the entered keywords may be incomplete or the language may be non-standard. This can lead to missed detections in keyword-matching-only search methods, thus affecting users' judgment of relevant laws.

[0004] Therefore, improving the comprehensiveness of legal data searches has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide an efficient legal data search method based on artificial intelligence to address the problem of how to improve the comprehensiveness of legal data searches.

[0006] This invention provides an efficient legal data search method based on artificial intelligence, which includes the following steps:

[0007] Obtain at least two target keywords input by the target user, record all keywords in all legal documents to be searched other than the target keywords as reference keywords, combine all target keywords to obtain at least one set of target keywords, and obtain new target keywords and their corresponding new reference keywords based on the occurrence of reference keywords in all legal documents to be searched that contain different arrangements of each set of target keywords.

[0008] All new target keywords and new reference keywords are recorded as new keywords. For any new keyword, the relevance between any new keyword and each other keyword is obtained based on the distance between any new keyword and other keywords in each legal document to be searched, and the number of times any new keyword appears in each legal document to be searched.

[0009] Obtain the relevance of each new target keyword and each new reference keyword; based on the relevance of each new target keyword and each new reference keyword, obtain the degree of association between each new target keyword and each new reference keyword.

[0010] Based on the frequency of occurrence of each new target keyword and each new reference keyword in each legal document to be retrieved, as well as the degree of association between each new target keyword and its corresponding new reference keyword, the recommendation level of each legal document to be retrieved is obtained, and search results are generated for the target user based on the recommendation level of each legal document to be retrieved.

[0011] Preferably, the step of obtaining new target keywords and their corresponding new reference keywords based on the occurrence of reference keywords in all legal documents to be retrieved containing different arrangements of each group of target keywords includes:

[0012] For any set of target keywords, perform a full permutation of the target keywords to obtain at least two combined keywords. If all the legal documents to be searched contain at most one of the combined keywords, then record all the target keywords in the target keywords as new target keywords, and use all the reference keywords as new reference keywords corresponding to each new target keyword.

[0013] If all legal documents to be retrieved contain at least two of the aforementioned combined keywords, then all legal documents to be retrieved containing the aforementioned combined keywords are recorded as target documents, and all reference keywords in the target documents are recorded as keywords to be analyzed.

[0014] Based on each of the combined keywords, all target documents are classified to obtain at least two document categories, with each of the combined keywords corresponding to one document category. Based on the frequency of occurrence of each keyword to be analyzed in each of the target documents in each document category, the content difference between all target documents corresponding to any set of target keywords is obtained.

[0015] Based on the content differences among all target documents corresponding to any set of target keywords, obtain new target keywords and their corresponding new reference keywords.

[0016] Preferably, the step of obtaining the content difference degree among all target documents corresponding to any set of target keywords based on the frequency of occurrence of each of the keywords to be analyzed in each of the target documents in each document category includes:

[0017] Based on the number of times each keyword to be analyzed appears in any target document, obtain the relative frequency of each keyword to be analyzed in any target document;

[0018] For any keyword to be analyzed, the relative frequency of the keyword to be analyzed in each of the target documents in any document category is obtained, and the average value of all the relative frequencies of the keyword to be analyzed in any document category is calculated to obtain the average importance of the keyword to be analyzed in any document category.

[0019] The average importance of any keyword to be analyzed in each document category is obtained. All document categories are paired to obtain at least one document category combination. The absolute value of the difference between the average importance of the two document categories in each document category combination is calculated. The absolute values ​​of the differences between all document category combinations are summed to obtain the importance difference of any keyword to be analyzed in all target documents.

[0020] The importance difference of each keyword to be analyzed in all target documents is obtained. The sum of all importance differences is linearly normalized to obtain the content difference between all target documents corresponding to any set of target keywords.

[0021] Preferably, obtaining new target keywords and their corresponding new reference keywords based on the content differences among all target documents corresponding to any set of target keywords includes:

[0022] If the content difference between all target documents corresponding to any set of target keywords is less than the preset content difference threshold, then all target keywords in any set of target keywords are recorded as new target keywords, and all reference keywords are used as new reference keywords corresponding to each new target keyword.

[0023] If the content difference between all target documents corresponding to any set of target keywords is greater than or equal to a preset content difference threshold, then all combination keywords existing in the legal documents to be searched are recorded as new target keywords, and any target keyword in any new target keyword is replaced with any reference keyword to obtain a new reference keyword corresponding to any new target keyword. The new reference keyword appears in at least one legal document to be searched.

[0024] Preferably, the step of obtaining the relevance between any new keyword and each other keyword based on the distance between any new keyword and other keywords in each legal document to be retrieved, and the number of times any new keyword appears in each legal document to be retrieved, includes:

[0025] Record all legal documents containing any of the new keywords as documents to be analyzed, and obtain at least one position of any of the new keywords in each document to be analyzed;

[0026] All reference keywords and target keywords other than the new keyword are denoted as other keywords. For any other keyword, if the other keyword does not exist in any document to be analyzed, the statistical correlation between the new keyword and the other keyword in the document to be analyzed is recorded as 0.

[0027] If any other keyword exists in any document to be analyzed, then at least one position of any other keyword is obtained in any document to be analyzed, and the statistical correlation between any new keyword and any other keyword in any document to be analyzed is obtained based on the number of characters between each position of any new keyword and each position of any other keyword in any document to be analyzed.

[0028] Based on the number of times any new keyword appears in any document to be analyzed, obtain the relative frequency of any new keyword in any document to be analyzed;

[0029] By combining the relative frequency of any new keyword in each document to be analyzed, and the statistical correlation between any new keyword and any other keyword in each document to be analyzed, the degree of relevance between any new keyword and any other keyword is obtained.

[0030] Preferably, obtaining the statistical correlation between any new keyword and any other keyword in any document to be analyzed based on the number of characters between each position of any new keyword and each position of any other keyword in any document to be analyzed includes:

[0031] In any document to be analyzed, for any position of any new keyword, obtain the number of characters between any position and each position of any other keyword, and record the minimum value among all the character counts as the minimum distance between any position and any other keyword;

[0032] Obtain the minimum distance between each position of any new keyword and any other keyword, and linearly normalize the reciprocal of the average of all minimum distances to obtain the statistical correlation between any new keyword and any other keyword in any document to be analyzed.

[0033] Preferably, the step of combining the relative frequency of any new keyword in each document to be analyzed, and the statistical correlation between any new keyword and any other keyword in each document to be analyzed, to obtain the relevance between any new keyword and any other keyword includes:

[0034] For any document to be analyzed, the relative frequency of any new keyword in the document to be analyzed is linearly normalized to obtain a relative frequency normalized value. The relative frequency normalized value is multiplied by the statistical correlation between the new keyword and any other keyword to obtain the degree of correlation between the new keyword and any other keyword in the document to be analyzed.

[0035] Calculate the mean of the relevance between any new keyword and any other keyword in all documents to be analyzed, and obtain the degree of relevance between any new keyword and any other keyword.

[0036] Preferably, the step of obtaining the correlation degree between each new target keyword and each corresponding new reference keyword based on the relevance degree between each new target keyword and each corresponding new reference keyword includes:

[0037] For any new target keyword and any corresponding new reference keyword, other keywords that belong to both the new target keyword and the new reference keyword are recorded as related keywords. The degree of relevance between the new target keyword and any related keyword is recorded as the first degree of relevance. The degree of relevance between the new reference keyword and any related keyword is recorded as the second degree of relevance. The absolute value of the difference between the first degree of relevance and the second degree of relevance is calculated to obtain the degree of relevance difference value of any related keyword.

[0038] The relevance difference value of each of the relevant keywords is obtained, and all the relevance difference values ​​are accumulated to obtain the accumulated result. The sum and reciprocal of the sum of the accumulated result and the preset constant are linearly normalized to obtain the degree of association between any new target keyword and any corresponding new reference keyword.

[0039] Preferably, the step of obtaining the recommendation level of each legal document to be retrieved based on the frequency of occurrence of each new target keyword in each legal document to be retrieved, and the degree of correlation between each new target keyword and its corresponding new reference keyword, includes:

[0040] For any legal document to be retrieved, obtain all new reference keywords corresponding to any new target keyword in the legal document to be retrieved, and obtain the relative frequency of the new target keyword and each of its corresponding new reference keywords based on the number of times the new target keyword and each of its corresponding new reference keywords appear in the legal document to be retrieved.

[0041] The degree of correlation between any new target keyword and each of its corresponding new reference keywords is used as the weight of the relative frequency of each of the corresponding new reference keywords. The relative frequencies of each of the corresponding new reference keywords are weighted and summed to obtain the weighted total frequency of all new reference keywords corresponding to any new target keyword.

[0042] The content relevance of any new target keyword is obtained by summing the relative frequency of any new target keyword with the total weighted frequency. The content relevance of all new target keywords in any legal document to be retrieved is then summed to obtain the recommendation level of any legal document to be retrieved.

[0043] Preferably, generating search results for the target user based on the recommendation level of each legal document to be retrieved includes:

[0044] Obtain the recommendation level of all legal documents to be retrieved, and sort them in descending order based on the recommendation level to obtain the search results for the target user.

[0045] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:

[0046] This invention determines whether there are significant differences in the content direction of legal documents containing different combinations of target keywords by analyzing the frequency of occurrence of reference keywords in the documents to be retrieved. This leads to the generation of new target keywords and their corresponding new reference keywords. The retrieval of legal documents is then performed based on these new target keywords and reference keywords, improving the comprehensiveness and accuracy of legal document searches. Furthermore, by identifying the frequency of occurrence of new target keywords in each legal document and considering the correlation between each new target keyword and its corresponding new reference keyword, the recommendation level of each legal document is determined. This generates search results for the target user, ensuring that the search results are based not only on the frequency of occurrence of the new target keywords in the legal documents but also on the frequency of occurrence of new reference keywords with strong correlation to the new target keywords. This effectively reduces the occurrence of missed detections and improves the comprehensiveness and accuracy of search results. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart of an efficient legal data search method based on artificial intelligence, provided in Embodiment 1 of the present invention. Detailed Implementation

[0049] Embodiments of this disclosure are described in detail below, with examples of these embodiments illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting it.

[0050] It should be noted that the terms "first," "second," etc., used in this disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.

[0051] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0052] See Figure 1 This is a flowchart of a method for efficient legal data search based on artificial intelligence, provided in Embodiment 1 of the present invention. Figure 1 As shown, the method may include:

[0053] Step S101: Obtain at least two target keywords input by the target user; record all keywords in all legal documents to be searched other than the target keywords as reference keywords; combine all target keywords to obtain at least one set of target keywords; and obtain new target keywords and their corresponding new reference keywords based on the occurrence of reference keywords in all legal documents to be searched that contain different arrangements of each set of target keywords.

[0054] When retrieving legal documents, existing keyword-matching search methods primarily retrieve all legal documents containing the target keywords entered by the user. However, considering that the content expressed by the target keywords is relatively limited—that is, there may be other related keywords—retrieving legal documents solely based on the target keywords may result in an incomplete list of retrieved documents.

[0055] Therefore, in this embodiment of the invention, all legal documents in the retrieval database are designated as legal documents to be retrieved, and a keyword dataset is generated using a pre-trained keyword extraction model (e.g., Legal-BERT). Any user is designated as the target user, and the keywords input by the target user are obtained, i.e., the keywords retrieved by the target user. In the keyword dataset, the keywords input by the target user are designated as target keywords, and keywords other than target keywords are designated as reference keywords. A correlation analysis is performed on each target keyword and each reference keyword to obtain reference keywords that are related to the target keywords. Legal documents are then retrieved based on the target keywords and keywords associated with them, improving the comprehensiveness of the legal document retrieval. The training of the keyword extraction model (e.g., Legal-BERT) is prior art and will not be elaborated upon here.

[0056] In legal documents, some keywords may appear as keyword phrases. The order of these keywords affects the resulting keyword phrases, and the content of legal documents containing these keyword phrases may be unrelated. That is, the content covered by these legal documents may differ significantly. For example, suppose a target user inputs "illegal possession of property." Extracting keywords from "illegal possession of property" yields target keywords such as "illegal," "possession," and "property." These target keywords could form keyword phrases like "illegal possession of property" or "possession of illegal property." Legal documents containing "illegal possession of property" might emphasize illegal possession, while those containing "possession of illegal property" might emphasize illegal property. Since the content of these two types of legal documents is dissimilar, legal documents containing different keyword phrases should be prioritized for the target user to improve the accuracy of search results. Simply presenting all legal documents containing "illegal," "possession," and "property" as search results may lead to decreased accuracy of search results, negatively impacting the user experience. Therefore, before conducting correlation analysis between each target keyword and each reference keyword, it is necessary to first consider different permutations and combinations of the target keywords to obtain new target keywords and their corresponding new reference keywords. This will allow for correlation analysis of each new target keyword and its corresponding new reference keywords, thereby improving the accuracy and comprehensiveness of the search results.

[0057] In this embodiment of the invention, assuming the number of target keywords is 'a', all target keywords are combined to obtain... Group target keywords: Based on the occurrence of reference keywords in all legal documents to be searched that contain different arrangements of each group of target keywords, new target keywords and their corresponding new reference keywords are obtained. Specifically:

[0058] Taking the target keywords of group f as an example, perform a full permutation of the target keywords of group f to obtain at least two combined keywords. For example, assuming that the target keywords of group f are D and E, the combined keywords are "DE" and "ED".

[0059] If all the legal documents to be retrieved contain at most one of the aforementioned combined keywords, it means that there is at most one legal combination among the multiple target keywords in the f-th group of target keywords in the legal documents to be retrieved. That is, there is no situation where the multiple combination forms of the multiple target keywords in the f-th group of target keywords express different meanings. Therefore, all the target keywords in the f-th group of target keywords are recorded as new target keywords, and all the reference keywords are used as new reference keywords corresponding to each new target keyword.

[0060] If all the legal documents to be retrieved contain at least two of the aforementioned combined keywords, it means that there are at least two different combinations of the multiple target keywords in the f-th group of target keywords. In this case, all the legal documents to be retrieved containing the aforementioned combined keywords are recorded as target documents, and all the reference keywords in the target documents are recorded as keywords to be analyzed.

[0061] Based on each of the combined keywords, all target documents are classified to obtain at least two document categories. Each of the combined keywords corresponds to one document category, that is, each document category represents one combined keyword. For example, assuming that the target keywords of the f-th group are D and E, the combined keywords are "DE" and "ED". All target documents containing "DE" are grouped into one category, denoted as category one, and all target documents containing "ED" are grouped into one category, denoted as category two. If a target document contains both "DE" and "ED", then the target document belongs to both category one and category two.

[0062] Based on the frequency of each keyword to be analyzed in each target document within each document category, the differences in document content between each document category are analyzed to obtain the new target keywords corresponding to the f-th group of target keywords and their corresponding new reference keywords. The implementation steps are as follows:

[0063] (1) Based on the number of times each keyword to be analyzed appears in each target document in each document category, the content difference between all target documents corresponding to the target keyword of group f is obtained.

[0064] Specifically: Taking the b-th keyword to be analyzed and the w-th target document as an example, MapReduce is used to count the number of times the b-th keyword to be analyzed appears in the w-th target document, and the relative frequency of the b-th keyword to be analyzed in the w-th target document is obtained, denoted as . ,Right now ,in, This represents the number of times the b-th keyword to be analyzed appears in the w-th target document. This represents the total number of words in the w-th target document, which is the sum of the occurrences of each keyword in the w-th target document. MapReduce is an existing technology and will not be elaborated on here.

[0065] Similarly, obtain the relative frequency of the b-th keyword to be analyzed in each target document in any document category, calculate the average value of all relative frequencies corresponding to the b-th keyword to be analyzed in any document category, and obtain the average importance of the b-th keyword to be analyzed in any document category.

[0066] Obtain the average importance of the b-th keyword to be analyzed in each document category, combine all document categories in pairs to obtain at least one document category combination, calculate the absolute value of the difference between the average importance of the two document categories in each document category combination, and sum the absolute values ​​of the differences between all document category combinations to obtain the importance difference of the b-th keyword to be analyzed in all target documents;

[0067] Obtain the importance difference of each keyword to be analyzed in all target documents, and linearly normalize the sum of all importance differences to obtain the content difference among all target documents corresponding to the f-th group of target keywords.

[0068] In one implementation, the formula for calculating the content difference between all target documents corresponding to the f-th group of target keywords is:

[0069]

[0070] in, This represents the content difference among all target documents corresponding to the target keywords in group f. This represents the number of keywords to be analyzed in all target documents corresponding to the target keywords in group f. This represents the number of all document category combinations corresponding to the target keyword group f. This represents the average importance of the b-th keyword to be analyzed within the first document category of the u-th document category combination. This represents the average importance of the b-th keyword to be analyzed in the second document category within the u-th document category combination. Represents the absolute value symbol. This represents the linear normalization function.

[0071] It should be noted that, The larger the value, the greater the difference in document content between the two document categories. Since each document category represents each combination of keywords corresponding to the f-th group of target keywords, a greater difference in the content of documents between the two document categories indicates a significant difference in the content direction between the legal documents to be retrieved that contain different combinations of keywords corresponding to the f-th group of target keywords. The larger.

[0072] (2) Based on the content difference between all target documents corresponding to the target keywords of group f, obtain the new target keywords and their corresponding new reference keywords corresponding to the target keywords of group f.

[0073] Specifically: Set the preset content difference threshold to 0.5. There is no restriction here. Implementers can set it according to the specific scenario. If the content difference between all target documents corresponding to the target keywords of group f is less than 0.5, it means that the document content direction between the legal documents to be searched that contain different combinations of keywords corresponding to the target keywords of group f is similar. Therefore, when searching for legal documents, there is no need to distinguish the different combinations of target keywords of group f. All target keywords in group f are recorded as new target keywords, and all reference keywords are used as new reference keywords corresponding to each new target keyword.

[0074] If the content difference between all target documents corresponding to the target keywords of group f is greater than or equal to 0.5, it indicates that there is a significant difference in the content direction of the legal documents to be retrieved that contain different combinations of keywords corresponding to the target keywords of group f. In this case, it is necessary to distinguish the different combinations of the target keywords of group f. That is, in the subsequent analysis, it is necessary to treat the different combinations of multiple target keywords in group f as a whole for analysis. Therefore, all the combinations of keywords corresponding to the target keywords of group f that exist in the legal documents to be retrieved are recorded as new target keywords. Since the new target keywords are combinations of at least two target keywords, in order to analyze the correlation between the target keywords and the reference keywords, the reference keywords must also be in combination form. That is, any target keyword in any new target keyword is replaced with any reference keyword to obtain the new reference keywords corresponding to the new target keywords. For example, assuming that the target keywords of group f are D and E, the new target keywords are "DE", and the reference keywords are F, then the new reference keywords are "DF" and "FE". The new reference keywords appear in at least one legal document to be retrieved.

[0075] Similarly, obtain all new target keywords and their corresponding new reference keywords.

[0076] Step S102: Record all new target keywords and new reference keywords as new keywords. For any new keyword, based on the distance between any new keyword and other keywords in each legal document to be searched, and the number of times the new keyword appears in each legal document to be searched, obtain the relevance between any new keyword and each other keyword.

[0077] Step S101 yields all new target keywords and their corresponding new reference keywords. Then, a correlation analysis is performed on the new target keywords and their corresponding new reference keywords to obtain new reference keywords that are related to the new target keywords. Legal documents are retrieved based on the new target keywords and the new reference keywords that are related to the new target keywords, thereby improving the comprehensiveness of the retrieval of legal documents.

[0078] In this embodiment of the invention, the relevance between each new target keyword and its corresponding reference keyword is determined based on the similarity of the degree of relevance between each new target keyword and each new reference keyword and each other keyword. For example, taking the y-th new target keyword and its corresponding s-th new reference keyword as an example, all reference keywords other than the y-th new target keyword and the target keyword are recorded as other keywords of the y-th new target keyword, and the degree of relevance between each other keyword of the y-th new target keyword and the y-th new target keyword is analyzed. Similarly, all reference keywords other than the s-th new reference keyword and the target keyword are recorded as other keywords of the s-th new reference keyword. The keyword analysis examines the relevance between each other keyword of the s-th new reference keyword and the s-th new reference keyword. Keywords that belong to both the y-th new target keyword and the s-th new reference keyword are denoted as related keywords of the y-th new target keyword and the s-th new reference keyword. The relevance between each related keyword and the y-th new target keyword is denoted as the first degree of relevance, and the relevance between each related keyword and the s-th new reference keyword is denoted as the second degree of relevance. If the first degree of relevance and the second degree of relevance for each related keyword are similar, it indicates a strong correlation between the y-th new target keyword and the s-th new reference keyword.

[0079] Based on the above analysis, a correlation analysis is performed on the new target keywords and their corresponding new reference keywords. For ease of explanation later, the new target keywords and their corresponding new reference keywords are denoted as new keywords. Taking the Ath new keyword as an example, the relevance between the Ath new keyword and each of the other keywords is first determined.

[0080] All reference keywords and target keywords, except for the Ath new keyword, are denoted as "other keywords." Taking the tth other keyword as an example, the relevance between the Ath new keyword and the tth other keyword is obtained based on the distance between the Ath new keyword and the tth other keyword in each legal document to be searched, and the number of times the Ath new keyword appears in each legal document to be searched. The specific steps are as follows:

[0081] (1) Record all legal documents containing the Ath new keyword as documents to be analyzed. Taking the sth document to be analyzed as an example, obtain the statistical correlation between the Ath new keyword and the tth other keyword in the sth document to be analyzed based on the distance between the Ath new keyword and each other keyword in the sth document to be analyzed.

[0082] Specifically, if the s-th document to be analyzed does not contain the t-th other keyword, then the statistical correlation between the A-th new keyword and the t-th other keyword in the s-th document to be analyzed is recorded as 0;

[0083] If the t-th other keyword exists in the s-th document to be analyzed, then obtain at least one position of the t-th other keyword in the s-th document to be analyzed;

[0084] In the s-th document to be analyzed, for any position of the A-th new keyword, obtain the number of characters between that position and each position of the t-th other keyword, and record the minimum value among all the character counts as the minimum distance between that position and the t-th other keyword;

[0085] Obtain the minimum distance between each position of the Ath new keyword and the tth other keyword, and linearly normalize the reciprocal of the average of all minimum distances to obtain the statistical correlation between the Ath new keyword and the tth other keyword in the sth document to be analyzed.

[0086] In one embodiment, if the s-th document to be analyzed contains the t-th other keyword, then the formula for calculating the statistical correlation between the A-th new keyword and the t-th other keyword in the s-th document to be analyzed is:

[0087]

[0088] in, This represents the statistical correlation between the A-th new keyword and the t-th other keyword in the s-th document to be analyzed. Let represent the average of the minimum distances between each position of the A-th new keyword and the t-th other keyword. This represents the linear normalization function.

[0089] It should be noted that, The smaller the value, the stronger the correlation between the A-th new keyword and the t-th other keyword in the s-th document being analyzed. The larger.

[0090] Similarly, obtain the statistical correlation between the Ath new keyword and the tth other keyword in each document to be analyzed.

[0091] (2) Obtain the relative frequency of the Ath new keyword in each document to be analyzed, and combine the relative frequency of the Ath new keyword in each document to be analyzed with the statistical correlation between the Ath new keyword and the tth other keyword in each document to be analyzed to obtain the degree of correlation between the Ath new keyword and the tth other keyword.

[0092] Specifically: For any document to be analyzed, the relative frequency of the Ath new keyword in the document to be analyzed is linearly normalized to obtain a normalized relative frequency value. The normalized relative frequency value is multiplied by the statistical correlation between the Ath new keyword and the tth other keyword to obtain the degree of correlation between the Ath new keyword and the tth other keyword in the document to be analyzed.

[0093] Calculate the mean relevance between the Ath new keyword and the tth other keyword in all documents to be analyzed, and obtain the relevance between the Ath new keyword and the tth other keyword.

[0094] In one implementation, the formula for calculating the relevance between the Ath new keyword and the tth other keyword is:

[0095]

[0096] in, This indicates the degree of relevance between the A-th new keyword and the t-th other keyword. This indicates the total number of documents to be analyzed. A document to be analyzed is a legal document containing the A-th new keyword that is to be retrieved. This represents the statistical correlation between the A-th new keyword and the t-th other keyword in the s-th document to be analyzed. This represents the normalized relative frequency value of the A-th new keyword in the s-th document to be analyzed.

[0097] It should be noted that, The larger the value, the stronger the relevance between the A-th new keyword and the t-th other keyword, and thus... The larger, The larger the value, the more frequently the Ath new keyword appears in the sth document being analyzed, and the stronger the relevance between the content of the sth document and the Ath new keyword. The higher the credibility, the better the s-th document to be analyzed can demonstrate the correlation between the A-th new keyword and the t-th other keyword.

[0098] Thus, we have obtained the relevance between the Ath new keyword and the tth other keyword. Similarly, we can obtain the relevance between the Ath new keyword and each other keyword.

[0099] Step S103: Obtain the relevance of each new target keyword and each new reference keyword. Based on the relevance of each new target keyword and each new reference keyword, obtain the degree of association between each new target keyword and each new reference keyword.

[0100] Following the method for obtaining the relevance between the Ath new keyword and each other keyword in step S102, obtain the relevance between each new keyword and each other keyword, that is, obtain the relevance between each new target keyword and each other keyword, as well as the relevance between each new reference keyword and each other keyword.

[0101] Furthermore, based on the relevance of each new target keyword to its corresponding new reference keyword, a correlation analysis is performed on each new target keyword and its corresponding new reference keyword to obtain the degree of correlation between each new target keyword and its corresponding new reference keyword. Specifically:

[0102] For any new target keyword and any corresponding new reference keyword, other keywords that belong to both the new target keyword and the new reference keyword are recorded as related keywords. The degree of relevance between the new target keyword and any related keyword is recorded as the first degree of relevance. The degree of relevance between the new reference keyword and any related keyword is recorded as the second degree of relevance. The absolute value of the difference between the first degree of relevance and the second degree of relevance is calculated to obtain the degree of relevance difference value of any related keyword.

[0103] The relevance difference value of each of the relevant keywords is obtained, and all the relevance difference values ​​are accumulated to obtain the accumulated result. The sum and reciprocal of the sum of the accumulated result and the preset constant are linearly normalized to obtain the degree of association between any new target keyword and any corresponding new reference keyword.

[0104] In one implementation, taking the y-th new target keyword and its corresponding s-th new reference keyword as an example, the formula for calculating the degree of relevance between the y-th new target keyword and its corresponding s-th new reference keyword is as follows:

[0105]

[0106] in, This indicates the degree of relevance between the y-th new target keyword and its corresponding s-th new reference keyword. Indicates the number of related keywords. This represents the difference in relevance between the m-th related keywords. This represents a preset constant used to prevent the denominator from being zero. In this embodiment of the invention, it is set... There are no restrictions here; implementers can set them according to the specific scenario. Represents the absolute value symbol. This represents the linear normalization function.

[0107] It should be noted that, The smaller the value, the more similar the relevance of the y-th new target keyword and the s-th new reference keyword to each related keyword, and thus... The larger the value, the stronger the correlation between the y-th new target keyword and the s-th new reference keyword.

[0108] Similarly, obtain the degree of correlation between each new target keyword and each corresponding new reference keyword.

[0109] Step S104: Based on the frequency of occurrence of each new target keyword and each new reference keyword in each legal document to be retrieved, and the degree of association between each new target keyword and its corresponding new reference keyword, obtain the recommendation level of each legal document to be retrieved, and generate search results for the target user based on the recommendation level of each legal document to be retrieved.

[0110] After obtaining the correlation between each new target keyword and its corresponding new reference keyword in step S103, the recommendation level of each legal document to be retrieved is obtained based on the frequency of occurrence of each new target keyword and each new reference keyword in each legal document to be retrieved, as well as the correlation between each new target keyword and its corresponding new reference keyword. Specifically:

[0111] For any legal document to be retrieved, obtain all new reference keywords corresponding to any new target keyword in the legal document to be retrieved, and obtain the relative frequency of the new target keyword and each of its corresponding new reference keywords based on the number of times the new target keyword and each of its corresponding new reference keywords appear in the legal document to be retrieved.

[0112] The degree of correlation between any new target keyword and each of its corresponding new reference keywords is used as the weight of the relative frequency of each of the corresponding new reference keywords. The relative frequencies of each of the corresponding new reference keywords are weighted and summed to obtain the weighted total frequency of all new reference keywords corresponding to any new target keyword.

[0113] The content relevance of any new target keyword is obtained by summing the relative frequency of any new target keyword with the total weighted frequency. The content relevance of all new target keywords in any legal document to be retrieved is then summed to obtain the recommendation level of any legal document to be retrieved.

[0114] In one implementation, taking the h-th legal document to be retrieved as an example, the formula for calculating the recommendation level of the h-th legal document to be retrieved is as follows:

[0115]

[0116] in, This indicates the recommendation level of the h-th legal document to be retrieved. This represents the number of all new target keywords in the h-th legal document to be retrieved. This represents the number of new reference keywords corresponding to the y-th new target keyword in the h-th legal document to be retrieved. This represents the relative frequency of the y-th new target keyword in the h-th legal document to be retrieved. This represents the relative frequency of the s-th new reference keyword for the y-th new target keyword in the h-th legal document to be retrieved. This indicates the degree of correlation between the y-th new target keyword and its corresponding s-th new reference keyword.

[0117] It should be noted that, The larger, The larger, The larger the value, the more frequently the y-th new target keyword appears in the h-th legal document to be retrieved, the stronger the correlation between the s-th new reference keyword and the y-th new target keyword, and the more frequently the s-th new reference keyword appears, the more the h-th legal document to be retrieved matches the target user's search content, and thus... The larger the value, the higher the recommendation level of the h-th legal document to be retrieved.

[0118] Similarly, the recommendation level of all legal documents to be retrieved is obtained, and all legal documents to be retrieved are sorted in descending order according to the recommendation level of each legal document. If multiple legal documents to be retrieved have the same recommendation level, they are sorted according to their publication time from most recent to oldest to obtain the search results for the target user. The search results are based not only on the frequency of the new target keyword in the legal documents to be retrieved, but also on the frequency of the frequency of the new reference keywords that are strongly related to the new target keyword in the legal documents to be retrieved, which effectively reduces the occurrence of missed detections and improves the comprehensiveness and accuracy of the search results.

[0119] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for efficient legal data search based on artificial intelligence, characterized in that, The aforementioned efficient legal data search method based on artificial intelligence includes: Obtain at least two target keywords input by the target user, record all keywords in all legal documents to be searched other than the target keywords as reference keywords, combine all target keywords to obtain at least one set of target keywords, and obtain new target keywords and their corresponding new reference keywords based on the occurrence of reference keywords in all legal documents to be searched that contain different arrangements of each set of target keywords. All new target keywords and new reference keywords are recorded as new keywords. For any new keyword, the relevance between any new keyword and each other keyword is obtained based on the distance between any new keyword and other keywords in each legal document to be searched, and the number of times any new keyword appears in each legal document to be searched. Obtain the relevance of each new target keyword and each new reference keyword; based on the relevance of each new target keyword and each new reference keyword, obtain the degree of association between each new target keyword and each new reference keyword. Based on the frequency of occurrence of each new target keyword and each new reference keyword in each legal document to be retrieved, as well as the degree of association between each new target keyword and its corresponding new reference keyword, the recommendation level of each legal document to be retrieved is obtained, and search results are generated for the target user based on the recommendation level of each legal document to be retrieved.

2. The method for efficient legal data search based on artificial intelligence according to claim 1, characterized in that, The step of obtaining new target keywords and their corresponding new reference keywords based on the occurrence of reference keywords in all legal documents to be retrieved that contain different arrangements of each group of target keywords includes: For any set of target keywords, perform a full permutation of the target keywords to obtain at least two combined keywords. If all the legal documents to be searched contain at most one of the combined keywords, then record all the target keywords in the target keywords as new target keywords, and use all the reference keywords as new reference keywords corresponding to each new target keyword. If all legal documents to be retrieved contain at least two of the aforementioned combined keywords, then all legal documents to be retrieved containing the aforementioned combined keywords are recorded as target documents, and all reference keywords in the target documents are recorded as keywords to be analyzed. Based on each of the combined keywords, all target documents are classified to obtain at least two document categories, with each of the combined keywords corresponding to one document category. Based on the frequency of occurrence of each keyword to be analyzed in each of the target documents in each document category, the content difference between all target documents corresponding to any set of target keywords is obtained. Based on the content differences among all target documents corresponding to any set of target keywords, obtain new target keywords and their corresponding new reference keywords.

3. The method for efficient legal data search based on artificial intelligence according to claim 2, characterized in that, The step of obtaining the content difference degree among all target documents corresponding to any set of target keywords based on the frequency of occurrence of each of the keywords to be analyzed in each of the target documents in each document category includes: Based on the number of times each keyword to be analyzed appears in any target document, obtain the relative frequency of each keyword to be analyzed in any target document; For any keyword to be analyzed, the relative frequency of the keyword to be analyzed in each of the target documents in any document category is obtained, and the average value of all the relative frequencies of the keyword to be analyzed in any document category is calculated to obtain the average importance of the keyword to be analyzed in any document category. The average importance of any keyword to be analyzed in each document category is obtained. All document categories are paired to obtain at least one document category combination. The absolute value of the difference between the average importance of the two document categories in each document category combination is calculated. The absolute values ​​of the differences between all document category combinations are summed to obtain the importance difference of any keyword to be analyzed in all target documents. The importance difference of each keyword to be analyzed in all target documents is obtained. The sum of all importance differences is linearly normalized to obtain the content difference between all target documents corresponding to any set of target keywords.

4. The efficient legal data search method based on artificial intelligence according to claim 2, characterized in that, The step of obtaining new target keywords and their corresponding new reference keywords based on the content difference between all target documents corresponding to any set of target keywords includes: If the content difference between all target documents corresponding to any set of target keywords is less than the preset content difference threshold, then all target keywords in any set of target keywords are recorded as new target keywords, and all reference keywords are used as new reference keywords corresponding to each new target keyword. If the content difference between all target documents corresponding to any set of target keywords is greater than or equal to a preset content difference threshold, then all combination keywords existing in the legal documents to be searched are recorded as new target keywords, and any target keyword in any new target keyword is replaced with any reference keyword to obtain a new reference keyword corresponding to any new target keyword. The new reference keyword appears in at least one legal document to be searched.

5. The efficient legal data search method based on artificial intelligence according to claim 1, characterized in that, The step of obtaining the relevance between any new keyword and each other keyword based on the distance between any new keyword and other keywords in each legal document to be retrieved, and the number of times any new keyword appears in each legal document to be retrieved, includes: Record all legal documents containing any of the new keywords as documents to be analyzed, and obtain at least one position of any of the new keywords in each document to be analyzed; All reference keywords and target keywords other than the new keyword are denoted as other keywords. For any other keyword, if the other keyword does not exist in any document to be analyzed, the statistical correlation between the new keyword and the other keyword in the document to be analyzed is recorded as 0. If any other keyword exists in any document to be analyzed, then at least one position of any other keyword is obtained in any document to be analyzed, and the statistical correlation between any new keyword and any other keyword in any document to be analyzed is obtained based on the number of characters between each position of any new keyword and each position of any other keyword in any document to be analyzed. Based on the number of times any new keyword appears in any document to be analyzed, obtain the relative frequency of any new keyword in any document to be analyzed; By combining the relative frequency of any new keyword in each document to be analyzed, and the statistical correlation between any new keyword and any other keyword in each document to be analyzed, the degree of relevance between any new keyword and any other keyword is obtained.

6. The efficient legal data search method based on artificial intelligence according to claim 5, characterized in that, The step of obtaining the statistical correlation between any new keyword and any other keyword in any document to be analyzed based on the number of characters between each position of any new keyword and each position of any other keyword in any document to be analyzed includes: In any document to be analyzed, for any position of any new keyword, obtain the number of characters between any position and each position of any other keyword, and record the minimum value among all the character counts as the minimum distance between any position and any other keyword; Obtain the minimum distance between each position of the new keyword and any other keyword, and linearly normalize the reciprocal of the average of all minimum distances to obtain the statistical correlation between the new keyword and any other keyword in any document to be analyzed.

7. The method for efficient legal data search based on artificial intelligence according to claim 5, characterized in that, The process of combining the relative frequency of any new keyword in each document to be analyzed, and the statistical correlation between any new keyword and any other keyword in each document to be analyzed, to obtain the degree of relevance between any new keyword and any other keyword, includes: For any document to be analyzed, the relative frequency of any new keyword in the document to be analyzed is linearly normalized to obtain a relative frequency normalized value. The relative frequency normalized value is multiplied by the statistical correlation between the new keyword and any other keyword to obtain the degree of correlation between the new keyword and any other keyword in the document to be analyzed. Calculate the mean of the relevance between any new keyword and any other keyword in all documents to be analyzed, and obtain the degree of relevance between any new keyword and any other keyword.

8. The method for efficient legal data search based on artificial intelligence according to claim 5, characterized in that, The step of obtaining the correlation degree between each new target keyword and each corresponding new reference keyword based on the relevance between each new target keyword and each corresponding new reference keyword includes: For any new target keyword and any corresponding new reference keyword, other keywords that belong to both the new target keyword and the new reference keyword are recorded as related keywords. The degree of relevance between the new target keyword and any related keyword is recorded as the first degree of relevance. The degree of relevance between the new reference keyword and any related keyword is recorded as the second degree of relevance. The absolute value of the difference between the first degree of relevance and the second degree of relevance is calculated to obtain the degree of relevance difference value of any related keyword. The relevance difference value of each of the relevant keywords is obtained, and all the relevance difference values ​​are accumulated to obtain the accumulated result. The sum and reciprocal of the sum of the accumulated result and the preset constant are linearly normalized to obtain the degree of association between any new target keyword and any corresponding new reference keyword.

9. The efficient legal data search method based on artificial intelligence according to claim 1, characterized in that, The process of obtaining the recommendation level for each legal document to be retrieved based on the frequency of occurrence of each new target keyword in each legal document to be retrieved, and the degree of correlation between each new target keyword and its corresponding new reference keyword, includes: For any legal document to be retrieved, obtain all new reference keywords corresponding to any new target keyword in the legal document to be retrieved, and obtain the relative frequency of the new target keyword and each of its corresponding new reference keywords based on the number of times the new target keyword and each of its corresponding new reference keywords appear in the legal document to be retrieved. The degree of correlation between any new target keyword and each of its corresponding new reference keywords is used as the weight of the relative frequency of each of the corresponding new reference keywords. The relative frequencies of each of the corresponding new reference keywords are weighted and summed to obtain the weighted total frequency of all new reference keywords corresponding to any new target keyword. The content relevance of any new target keyword is obtained by summing the relative frequency of any new target keyword with the total weighted frequency. The content relevance of all new target keywords in any legal document to be retrieved is then summed to obtain the recommendation level of any legal document to be retrieved.

10. The method for efficient legal data search based on artificial intelligence according to claim 1, characterized in that, The process of generating search results for the target user based on the recommendation level of each legal document to be searched includes: Obtain the recommendation level of all legal documents to be retrieved, and sort them in descending order based on the recommendation level to obtain the search results for the target user.