A method and system for retrieving legal provisions

By calculating the classification information, frequency of occurrence and correlation of legal provisions, weakening the search results and prioritization sorting, the problems of inefficient search efficiency and insufficient accuracy of traditional Boolean models when dealing with large-scale and multi-field legal norms are solved, and more efficient and accurate legal provision retrieval is achieved.

CN119884281BActive Publication Date: 2025-05-30GUANGDONG BOWEI CHUANGYUAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361576.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-30
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When traditional Boolean models deal with large-scale, multi-field and multi-level legal norms, they have low retrieval efficiency and insufficient results accuracy, and users need to manually filter a large amount of irrelevant information.

Method used

By obtaining the set of relevant entries in the Boolean model, the classification information, occurrence frequency, correlation probability and correlation degree of the entries are calculated, and the initial search results are weakened and the priority sorting of the final search results are performed to obtain the final displayed results.

Benefits of technology

It significantly improves the relevance and accuracy of search results, reduces the time and energy of users to screen the information required in a large number of results, and improves user experience and retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884281B_ABST
    Figure CN119884281B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of electronic digital data processing, and particularly to a legal provision retrieval method and system. The method includes: obtaining a set of relevant terms in a Boolean model, and retrieving in a database to obtain a preliminary retrieval result and a final retrieval result; according to the classification information of the terms, counting the occurrence frequencies of each term in different classifications, calculating the relevant probabilities in each classification, and selecting the maximum relevant probability as the importance weight value of the term; calculating the correlation degree between any two terms in the set of relevant terms, weakening the term combinations with higher occurrence frequencies, and accumulating the weakening degree to obtain the overall weakening degree; performing weighted summation on the importance weight value and the occurrence frequency of the term in the legal provisions, and correcting according to the overall weakening degree to obtain the priority degree of the final retrieval result. The present invention improves the relevance and accuracy of the retrieval result by performing weakening processing and priority degree sorting on the retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic digital data processing. In particular, it relates to a legal provision retrieval method and system. Background Art

[0002] Legal provision retrieval refers to the process of quickly and accurately finding legal provisions or information related to user needs in a vast amount of legal literature (such as laws and regulations, judicial interpretations, cases, academic articles, etc.) through specific technologies or methods.

[0003] With the increasing complexity of the legal system and the continuous update of legal provisions, when encountering various problems in life and needing to maintain one's rights and interests through legal provisions, people usually search for relevant legal regulations on the Internet. The Internet has greatly improved the accessibility of legal information. People can quickly find relevant legal provisions, judicial interpretations, and cases through search engines, legal databases, government official websites, etc. Therefore, the Internet has made legal knowledge no longer limited to legal practitioners, and ordinary citizens can also understand their rights and obligations through online searches, thereby better safeguarding their legitimate rights and interests.

[0004] Currently, legal issues often involve legal provisions in multiple fields. Moreover, in addition to national laws, there are also multi-level legal norms such as local regulations and judicial interpretations. These legal norms complement each other, but at the same time, they also increase the complexity of retrieval. With the complexity of the legal system and the continuous update of legal provisions, traditional Boolean model retrieval methods face many challenges when dealing with large-scale, multi-field, and multi-level legal norms. During the process of legal provision retrieval using the conventional Boolean model, the retrieval results of each term are directly output, resulting in the user still having to find the desired text from a large number of search results after the query, reducing the accuracy and convenience of the Boolean model query. Summary of the Invention

[0005] To solve the problems of low retrieval efficiency, insufficient result accuracy, and the need for users to manually screen a large amount of irrelevant information when the traditional Boolean model processes large-scale, multi-field, and multi-level legal norms, the present invention provides solutions in the following aspects.

[0006] In a first aspect, a legal provision retrieval method includes: obtaining a set of relevant terms in a Boolean model, performing a database retrieval on the terms in the set of relevant terms to obtain a preliminary retrieval result that includes any term in the set of relevant terms, and screening out, from the preliminary retrieval result, a retrieval result that strictly matches the query condition of the Boolean expression as the final retrieval result; obtaining the classification information of the terms, counting the occurrence frequencies of the terms in each classification information, calculating the correlation probabilities of the terms in each classification information, and selecting the maximum correlation probability in each classification information as the importance weight of the term, where the classification information refers to various classification attributes related to legal provisions; calculating the correlation degree between any two terms in the set of relevant terms, obtaining the term combinations with higher occurrence frequencies in the set of relevant terms, performing a weakening process based on the correlation degree, obtaining the weakening degree of each term combination and summing them up to obtain the overall weakening degree of the preliminary retrieval result; performing a weighted sum of the importance weight and the occurrence frequency of the term in the legal provision, and making a correction based on the overall weakening degree to obtain the priority degree of the final retrieval result, sorting the final retrieval result according to the priority degree, and obtaining the final display result.

[0007] By comprehensively considering the classification information, occurrence frequency, correlation probability of the terms, and the correlation degree between the terms, the matching degree between the legal provision and the user's query intention is accurately evaluated, thereby significantly improving the relevance and accuracy of the retrieval result. At the same time, by performing a weighted sum and correction on the retrieval result based on the importance weight, occurrence frequency, and overall weakening degree of the term, the sorting of the retrieval result is optimized, avoiding the inflated priority degree caused by repeated weighting of the terms, and making the result sorting more reasonable. Through precise screening and sorting, the time and effort required for the user to screen the required information from a large number of results are reduced, significantly improving the user experience and retrieval efficiency. The final display result is presented in the form of a list sorted by priority degree, enabling the user to quickly locate the most relevant legal provision and meeting the user's requirements for efficiency and convenience in legal retrieval.

[0008] Preferably, the preliminary retrieval result is all legal provisions that contain relevant terms retrieved based on the Boolean model, and the final retrieval result is the legal provisions that completely meet the user's retrieval conditions screened out from the preliminary retrieval result;

[0009] In response to a legal provision containing any term in the Boolean model, the legal provision is determined as the preliminary retrieval result, and in response to the legal provision containing all terms in the Boolean model, the legal provision is considered as the final detection result.

[0010] By clearly distinguishing the preliminary search results from the final search results, accurate positioning and efficient screening of the search results are achieved. The preliminary search results are determined based on the presence of any term in the Boolean model, ensuring a comprehensive search scope without missing any potentially relevant articles; while the final search results strictly screen out the articles that meet all the terms of the Boolean model, precisely matching the user's search conditions and improving the pertinence and accuracy of the search results.

[0011] Preferably, the relevant probability includes:

[0012] Taking any term as the target term, calculate the ratio between the sum of the occurrence frequencies of the target term in all legal articles in a specific classification information and the sum of the occurrence frequencies of the target term in all legal articles in all classification information, to obtain the relative importance of the target term in the specific classification information;

[0013] Calculate the proportion of the number of legal articles of the target term in the specific classification information to the total number of all search results, to obtain the concentration degree of the target term in the specific classification information;

[0014] Take the product between the relative importance and the concentration degree as the relevant probability of the target term in the specific classification information.

[0015] By calculating the relative importance and concentration degree of the target term in a specific classification and taking the product as the relevant probability, the importance and relevance of the term in different classifications can be accurately evaluated. The relative importance reflects the activity degree of the term in this classification, and the concentration degree reflects the concentrated distribution situation of the term in this classification, which can effectively distinguish the actual contributions of the term in different classifications, thereby improving the relevance and accuracy of the search results.

[0016] Preferably, the relevant probability further includes:

[0017] Taking any term as the target term, calculate the ratio between the sum of the occurrence frequencies of the target term in all legal articles in a specific classification information and the total number of the search results of the target term, and take the ratio as the relevant probability of the target term.

[0018] Preferably, the relevant degree includes:

[0019] Divide the number of times that any two terms appear simultaneously in the preliminary search results by the square root of the product of the total number of times each of the two terms appears in the preliminary search results, and take the square root result as the relevant degree of the two terms.

[0020] By quantifying the association strength between terms, the correlation between terms can be more accurately evaluated, thereby reasonably adjusting the weights of terms in the retrieval results, avoiding the inflated priority caused by the repeated weighting of terms, improving the relevance and accuracy of the retrieval results, and ensuring that users can more efficiently find the most relevant legal provisions.

[0021] Preferably, the degree of weakening includes:

[0022] Calculate the co-occurrence frequency of any two terms in the preliminary retrieval results. In response to the co-occurrence frequency being greater than the average occurrence frequency of all terms, weaken the correlation degree of the two terms; conversely, if it is less than or equal to the average occurrence frequency of all terms, no weakening is required;

[0023] Among them, the weakening method is: use the negative exponential function to weaken the correlation degree of the two terms to be used to reduce the priority of the retrieval results.

[0024] By calculating the co-occurrence frequency of any two terms in the preliminary retrieval results and comparing it with the average occurrence frequency of all terms, the correlation degree of the terms is dynamically adjusted. When the co-occurrence frequency is higher than the average value, use the negative exponential function to weaken the correlation degree of the two terms, thereby reducing the priority of the retrieval results; conversely, if the co-occurrence frequency is lower than or equal to the average value, the correlation degree remains unchanged. This dynamic adjustment mechanism can effectively avoid the inflated priority caused by the repeated weighting of terms and ensure the accuracy and relevance of the retrieval results.

[0025] Preferably, the overall degree of weakening satisfies the following relational expression:

[0026] ;

[0027] In the formula, represents the overall degree of weakening of the th preliminary retrieval result, represents the total number of terms in the relevant term set, represents the th preliminary retrieval result, the th term and the th term of the

[0028] Preferably, the priority satisfies the following relational expression:

[0029] ;

[0030] In the formula, represents the priority of the th preliminary retrieval result, represents the overall degree of weakening of the th preliminary retrieval result, represents the total number of entries in the relevant entry set, represents the importance weight of the th entry, represents the th entry's occurrence frequency in the final search results filtered from the th preliminary search result.

[0031] Preferably, the obtaining of the final display result includes:

[0032] Sorting the final search results in descending order based on the priority level and outputting them in a list form for viewing the most relevant legal provisions.

[0033] In a second aspect, a legal provision retrieval system includes: a processor and a memory, and the memory stores computer program instructions, which implement the above-mentioned legal provision retrieval method when executed by the processor.

[0034] The present invention has the following effects:

[0035] 1. By calculating the relevant probabilities of entries in each classification and the relevance degree between entries, and performing weakening processing on the preliminary search results and sorting the final search results according to the priority level, the present invention can effectively screen out the legal provisions that best match the user's query intention, improve the relevance and accuracy of the search results, and reduce the time and effort for the user to screen the required information from a large number of results.

[0036] 2. By optimizing the search algorithm, the present invention reduces the artificially inflated priority level caused by repeated weighting of entries, ensuring the quality and reliability of the search results. At the same time, sorting the final search results according to the priority level and outputting them in a list form enables the user to quickly view the most relevant legal provisions, significantly improving the user experience and search efficiency, and meeting the diverse needs of users in legal retrieval.

[0037] 3. By comprehensively considering various factors such as the classification information, occurrence frequency, and relevance degree of entries, the present invention can dynamically adjust the priority level of the search results, enhancing the adaptability of the retrieval system to complex legal systems and multi-domain legal norms, enabling it to better cope with the continuous update and change of legal provisions, and providing users with more intelligent and accurate retrieval services. Description of the Drawings

[0038] Figure 1 is a flowchart of the methods of steps S1 - S4 in a legal provision retrieval method according to an embodiment of the present invention.

[0039] Figure 2 is a structural block diagram of a legal provision retrieval system according to an embodiment of the present invention. Specific implementation manner

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments.

[0041] Refer to Figure 1 , a legal provision retrieval method includes steps S1 - S4, specifically as follows:

[0042] S1: Obtain the set of relevant terms in the Boolean model, perform database retrieval on the terms in the set of relevant terms, obtain the preliminary retrieval results including any term in the set of relevant terms, and screen out the retrieval results that strictly match the query conditions of the Boolean expression from the preliminary retrieval results as the final retrieval results.

[0043] It should be noted that the user enters keywords or phrases in the search box, and these input terms form a set of terms. The system automatically expands relevant synonyms, related terms, professional terms, etc. according to the terms entered by the user to form a more comprehensive set of relevant terms.

[0044] Exemplarily, the user enters keywords or phrases in the search box. For example, the user enters: "contract breach"; the system expands: the set of terms includes "contract breach", "liability for breach of contract", "contract rescission", "liquidated damages", etc.

[0045] The preliminary retrieval results are all legal provisions containing relevant terms retrieved based on the Boolean model, and the final retrieval results are the legal provisions that fully meet the user's retrieval conditions screened out from the preliminary retrieval results;

[0046] In response to a legal provision containing any term in the Boolean model, the legal provision is determined as a preliminary retrieval result. In response to a legal provision containing all terms in the Boolean model, the legal provision is considered as the final detection result.

[0047] Exemplarily, enter: 1 and 2 in the Boolean model. Statute 1: Contains the term "1", but does not contain the term "2", belongs to the preliminary retrieval result, but does not belong to the final retrieval result; Statute 2: Contains both the term "1" and the term "2", belongs to the preliminary retrieval result, but also belongs to the final retrieval result at the same time.

[0048] It should be noted that the difference between the preliminary retrieval results and the final retrieval results is: preliminary retrieval results: legal provisions containing any term in the Boolean model, used to expand the retrieval scope; final retrieval results: legal provisions containing all terms in the Boolean model, used to ensure the accuracy and relevance of the retrieval results.

[0049] For further explanation, legal provisions each have their own classification information. Specifically, legal provisions are subdivided based on legal fields. Exemplarily, the classification information includes: civil law, criminal law, commercial law, administrative law, and intellectual property law, etc. The total number of general classifications of legal provisions obtained is denoted as , and all retrieval results of each entry are obtained .

[0050] S2: Obtain the classification information of the entry, count the frequency of each entry in each classification information, calculate the correlation probability of the entry in each classification information, and select the largest correlation probability in each classification information as the importance weight value of the entry. Among them, the classification information refers to various classification attributes related to legal provisions.

[0051] The correlation probability includes:

[0052] Taking any entry as the target entry, calculate the ratio between the sum of the frequencies of the target entry in all legal provisions in a specific classification information and the sum of the frequencies of all legal provisions in all classification information, and obtain the relative importance of the target entry in the specific classification information;

[0053] Calculate the proportion of the number of legal provisions of the target entry in a specific classification information to the total number of retrieval results, and obtain the concentration degree of the target entry in the specific classification information;

[0054] Take the product between the relative importance and the concentration degree as the correlation probability of the target entry in the specific classification information.

[0055] Specifically, the correlation probability satisfies the following relational expression:

[0056] ;

[0057] In the formula, represents the correlation probability of the rd entry in all retrieval results of the th category, represents the number of retrieval results of the rd entry in the th category, represents the frequency of the th entry in the th legal provision in the th category, represents the total number of general classifications of legal provisions, represents the th total number of retrieval results of the entry.

[0058] That is to say, represents the entry in the The sum of the frequencies of occurrence in the classification search results reflects the in the total frequency of occurrence of the entry. Analyze the proportion of the frequency of occurrence in each category to the frequency of occurrence in all classification information. The higher the proportion, the greater the entry is related to the category. The legal provisions of the entry can be directly used to distinguish the classification.

[0059] Evaluate the relevance of an entry in classification information through the proportion of the global frequency of occurrence. If the frequency of occurrence of an entry in classification information is high and the frequency of occurrence in other classification information is low, the correlation probability will be high. Through reflect the concentration degree of the search results of the entry in classification information. If it is mainly concentrated in a certain classification information, the corresponding correlation probability will be higher, which is more suitable for evaluating the global relevance of an entry in a certain classification information and can better reflect the classification differences of an entry in different classification information.

[0060] In addition, in another embodiment, it further includes:

[0061] Taking any entry as the target entry, calculate the ratio between the sum of the frequencies of occurrence of the target entry in all legal provisions in specific classification information and the total number of search results of the target entry, and take the ratio as the correlation probability of the target entry.

[0062] Specifically, the correlation probability satisfies the following relational expression:

[0063] ;

[0064] In the formula, represents the th entry in the class, the correlation probability of all search results, represents the th entry in the class, the number of search results, represents the th entry in the class, the th legal provision, the frequency of occurrence, represents the th entry, the total number of search results.

[0065] reflects the th entry in the class, the average frequency of occurrence. That is to say, relative to the averaging process of the total number of search results, the higher the correlation probability, the th entry in the The frequency of occurrence in the class is significantly higher than the average frequency in all retrieval results, indicating that the th entry has stronger discrimination ability and importance in the class; conversely, the discrimination ability in the

[0066] class is weaker. By combining the frequency of occurrence of the th entry in the class with the total number of retrieval results of the th entry, that is, after normalization processing, it pays more attention to the absolute occurrence frequency of the entry in specific classification information, directly reflects the average occurrence frequency of the entry in specific classification information, does not depend on the distribution of the entry in other classification information, and is applicable to the situation where the number of classifications is large and the distribution is relatively uniform.

[0067] Specifically, the importance weight of the entry satisfies the following relationship:

[0068] ;

[0069] In the formula, represents the importance weight of the th entry, represents the maximum value function in all classification information, represents the total number of classifications of legal provisions, represents the th entry's class's relevant probability of all retrieval results.

[0070] That is to say, if an entry has a high relevant probability in a certain classification information, then this entry has a strong discrimination ability for this classification, so it should be given a higher weight. Therefore, the importance weight of the entry can better reflect the maximum discrimination ability of the entry in all classification information, providing a more accurate basis for subsequent sorting and retrieval.

[0071] S3: Calculate the correlation degree between any two entries in the relevant entry set, obtain the entry combinations with higher occurrence frequencies in the relevant entry set, perform weakening processing based on the correlation degree, obtain the weakening degree of each entry combination and sum them up to obtain the overall weakening degree of the preliminary retrieval result.

[0072] The correlation degree includes:

[0073] Divide the number of times any two entries appear simultaneously in the preliminary retrieval result by the sum of the total number of times the two entries appear in the preliminary retrieval result and take the square root, and use the square root as the correlation degree between the two entries.

[0074] Specifically, the degree of correlation satisfies the following relational expression:

[0075] ;

[0076] In the formula, represents the degree of correlation between the -th entry and the -th entry, represents the number of retrieval results that simultaneously contain the -th entry and the -th entry in the preliminary retrieval results, and respectively represent the total number of retrieval results of the -th entry and the -th entry.

[0077] It should be noted that in the process of calculating the priority by weighting with the importance weights of the entries, there are cases where the degree of correlation between entries is relatively high and the importance weights are also relatively high. For example, the entry has strong classification ability and a high importance weight, but the appearance of the entry often occurs together with the entry . Then, for the retrieval results where the entry and the entry appear simultaneously, when calculating the priority, they will be weighted repeatedly to a relatively large extent according to the entry and the entry respectively. However, in fact, the strong classification abilities of the two entries are the same, resulting in a falsely high priority. Therefore, the priority of such entries needs to be adjusted.

[0078] That is to say, only the entries with a relatively high appearance frequency in the preliminary retrieval results need to consider the weakening of the degree of correlation on the preliminary retrieval results. Therefore, the weakening degree is calculated according to the appearance frequency of each entry in the entries.

[0079] The weakening degree includes:

[0080] Calculate the joint appearance frequency of any two entries in the preliminary retrieval results. In response to the joint frequency being greater than the average appearance frequency of all entries, the degree of correlation between the two entries is weakened; conversely, if it is less than or equal to the average appearance frequency of all entries, no weakening is required;

[0081] Among them, the weakening method is: use a negative exponential function to weaken the degree of correlation between the two entries to be used for reducing the priority of the retrieval results.

[0082] Specifically, the weakening degree satisfies the following relational expression:

[0083] ;

[0084] In the formula, represents the weakening degree of the -th entry and the -th entry in the -th preliminary search result. represents the correlation degree between the -th entry and the -th entry. represents the occurrence frequency of the -th entry in the -th preliminary search result. represents the occurrence frequency of the -th entry in the -th preliminary search result. represents the average occurrence frequency of all entries in the -th preliminary search result. represents the exponential function with the natural number as the base.

[0085] That is to say, , indicating that the joint occurrence frequency of the -th entry and the -th entry in the -th preliminary search result is higher than the average value, and weakening is required according to their correlation degree. Otherwise, no weakening is needed. Therefore, is set to 1.

[0086] The overall weakening degree satisfies the following relational expression:

[0087] ;

[0088] In the formula, represents the overall weakening degree of the -th preliminary search result. represents the total number of entries in the set of relevant entries. represents the weakening degree of the -th entry and the -th entry in the -th preliminary search result.

[0089] It should be noted that the outer loop starts from , and the inner loop starts from , ensuring that all possible combinations of entries are considered, and each pair of entries is calculated only once; by adjusting the product result with the exponent , which is the reciprocal of the number of all entry combinations, the product result is normalized to obtain a reasonable average weakening degree.

[0090] S4: Weighted sum the importance weight and the frequency of occurrence of the entry in the legal provisions, and correct based on the overall weakening degree to obtain the priority of the final retrieval result, sort the final retrieval result according to the priority, and obtain the final display result.

[0091] Specifically, the priority satisfies the following relational expression:

[0092] ;

[0093] In the formula, represents the priority of the th legal provision, represents the overall weakening degree of the th preliminary retrieval result, represents the total number of entries in the relevant entry set, represents the th importance weight of the entry, represents the th occurrence frequency of the th entry in the

[0094] That is to say, the higher the priority , the stronger the relevance between the legal provision and the user's search intention. The higher the importance weight of the entry, and the higher its occurrence frequency in the legal provision, the higher the priority of the legal provision will be.

[0095] Sort the final retrieval results in descending order based on the priority and output them in a list form for viewing the most relevant legal provisions.

[0096] That is to say, sort all legal provisions from high to low according to the priority. In response to a higher priority, it indicates that the legal provision better matches the user's search intention and should be ranked in the front. On the contrary, the lower the priority, the weaker the relevance of the legal provision, and it should be ranked in the back.

[0097] It should be noted that sorting according to the priority of each retrieval result ensures that the user can quickly find the most relevant and most in-line legal provisions. The calculation of the priority is based on the importance weight of the entry and its occurrence frequency in the legal provision, so the sorting result can reflect the relevance and accuracy of the retrieval result.

[0098] The present invention also provides a legal provision retrieval system. As Figure 2As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a legal provision retrieval method according to the first aspect of the present invention. The system also includes a communication bus, a communication interface, and other components well-known to those skilled in the art. Their settings and functions are known in the art, so they will not be elaborated here.

[0099] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.

Claims

1. A legal text retrieval method, characterized in that: include: Obtaining a set of related terms in the Boolean model, performing a database search on the terms in the set of related terms, obtaining a preliminary search result containing any term in the set of related terms, and screening out a search result that strictly matches the Boolean expression query condition from the preliminary search result as a final search result; Obtaining classification information of the terms, counting the frequency of occurrence of each term in each classification information, calculating the relevant probability of the term in each classification information, and selecting the largest relevant probability in each classification information as the importance weight of the term, wherein the classification information refers to various classification attributes related to the legal provisions; Calculate the correlation between any two terms in the related term set, obtain the term combination with a higher frequency in the related term set, perform weakening processing based on the correlation, obtain the weakening degree of each term combination and sum them up to obtain the overall weakening degree of the preliminary search results; The degree of weakening includes: Calculate the joint occurrence frequency of any two terms in the preliminary search results. If the joint occurrence frequency is greater than the average occurrence frequency of all terms, the relevance of the two terms is weakened; otherwise, if the joint occurrence frequency is less than or equal to the average occurrence frequency of all terms, no weakening is required. The weakening method is: using a negative exponential function to weaken the relevance of two terms, so as to reduce the priority of the search results; The overall weakening degree satisfies the following relationship: ; In the formula, Indicates The overall weakening of the initial search results, Represents the total number of terms in the related term set, Indicates The first search result Entries and The degree of weakening of each entry; The importance weight and the frequency of occurrence of the term in the legal text are weighted and summed, and corrections are made based on the overall weakening degree to obtain the priority of the final search results. The final search results are sorted according to the priority to obtain the final display results.

2. A legal text retrieval method according to claim 1, characterized in that: The preliminary search results are all legal provisions containing relevant terms retrieved based on the Boolean model, and the final search results are legal provisions that fully meet the user's search conditions screened out from the preliminary search results; In response to the legal text containing any term in the Boolean model, the legal text is considered to be a preliminary search result. In response to the legal text containing all terms in the Boolean model, the legal text is considered to be a final detection result.

3. A legal text retrieval method according to claim 1, characterized in that: The relevant probabilities include: Taking any term as the target term, calculate the ratio between the sum of the frequency of occurrence of the target term in all legal articles in a specific classification and the sum of the frequency of occurrence of all legal articles in all classifications, and obtain the relative importance of the target term in the specific classification; Calculate the ratio of the number of legal articles of the target term in the specific classification information to the total number of search results to obtain the concentration of the target term in the specific classification information; The product of the relative importance and the concentration level is used as the relevance probability of the target term in the specific classification information.

4. A legal text retrieval method according to claim 1, characterized in that: The related probability also includes: Taking any term as the target term, the ratio between the sum of the occurrence frequencies of the target term in all legal provisions in the specific classification information and the total number of search results of the target term is calculated, and the ratio is taken as the relevant probability of the target term.

5. A legal text retrieval method according to claim 1, characterized in that: The degree of relevance includes: Divide the number of times any two terms appear simultaneously in the preliminary search results by the total number of times the two terms appear in the preliminary search results respectively, and take the square root of the result. The square root is used as the correlation degree between the two terms.

6. A legal text retrieval method according to claim 1, characterized in that: The priority satisfies the following relationship: ; In the formula, Indicates The priority of the initial search results, Indicates The overall weakening of the initial search results, Represents the total number of terms in the related term set, Indicates The importance weight of each term, Indicates The final search results are selected from the preliminary search results. The frequency of occurrence of a term.

7. A legal text retrieval method according to claim 1, characterized in that: The obtaining of the final display result includes: The final search results are sorted from largest to smallest based on priority and output in a list format for viewing the most relevant legal provisions.

8. A legal text retrieval system, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the legal text retrieval method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Legal database establishing method and legal retrieving service method

    CN104008171A

  • Legal provision retrieval method for legal data embedding optimization and retrieval effect evaluation

    CN119377384A