Cloud-computing-based legal service platform retrieval system
By combining initial screening and secondary screening, and utilizing a cloud computing platform to assess the timeliness and regionality of legal documents, the problem of low retrieval efficiency and poor system stability of existing legal service platforms is solved, achieving efficient and convenient legal document retrieval and system optimization and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2024-01-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing legal service platforms suffer from problems such as insufficient timeliness and regionality in document retrieval, disordered display of legal documents on the query page, and difficulty in timely monitoring, early warning, and system optimization and maintenance, resulting in low retrieval efficiency and poor targeting.
By combining initial screening and secondary screening, and through the collaborative work of the document retrieval database, information collection module, cloud computing platform and retrieval display module, legal documents can be searched, retrieved and sorted for display. Evaluation is performed based on content relevance, retrieval accuracy, update latency and page responsiveness, and anomaly signals or management signals are generated for system optimization and maintenance.
It improves the targeting and efficiency of legal document retrieval, ensures the continuous stability of the system and the convenience of user retrieval, and realizes refined monitoring, early warning and optimization maintenance of the retrieval system.
Smart Images

Figure CN117708275B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, and in particular to a cloud-based legal service platform retrieval system. Background Technology
[0002] A legal service platform is an online platform that provides legal services such as legal consultation, legal document preparation, and lawyer services. Users can find suitable lawyers, consult experts, or obtain relevant legal documents and materials through the legal service platform. When using the legal service platform to search for legal documents, users need to carefully identify and screen the search results to ensure that the documents obtained meet their own needs and requirements.
[0003] Existing legal service platforms suffer from insufficient timeliness and regionality when searching for documents. Due to the updating and iteration of legal provisions, the legal documents retrieved are outdated or not applicable to the current region, resulting in inaccurate searches and inapplicable documents. Furthermore, the legal documents displayed on the search page are disorganized, making it difficult for users to accurately find the most relevant documents. This leads to poor targeting of legal document searches, resulting in poor document search performance, affecting search efficiency, and consequently impacting the continuous stability of the system. It also makes it difficult to monitor, warn, and optimize the system in a timely manner.
[0004] To address the aforementioned technical shortcomings, a solution is proposed. Summary of the Invention
[0005] The purpose of this invention is to address the problems of insufficient timeliness and regionality in existing systems for retrieving documents, disordered display of legal documents on query pages, difficulty in timely monitoring, early warning, and system optimization and maintenance, and the resulting low retrieval efficiency and poor targeting. This invention achieves the retrieval and sorting of legal documents through a combination of initial and secondary screening, ensuring the timeliness and regionality of legal documents. Page sorting is performed by calculating content relevance, making user retrieval more efficient and convenient. By analyzing document and retrieval information through a cloud computing platform, the retrieval effect is comprehensively analyzed and evaluated, enabling monitoring, early warning, and optimization maintenance of the retrieval system, ensuring its continuous stability. Furthermore, through refined and in-depth retrieval, the targeting and efficiency of legal document retrieval are improved.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] The cloud-based legal service platform retrieval system includes a document retrieval database, an information collection module, a cloud computing platform, and a retrieval and display module, wherein the document retrieval database, the cloud computing platform, the information collection module, and the retrieval and display module are connected by signals.
[0008] The document retrieval database is used to store and update legal documents: legal documents are retrieved and stored through background uploads, time and region codes are set and marked, and legal documents are dynamically updated;
[0009] The information acquisition module is used to collect file information and retrieve search information: it obtains file information through a file retrieval database and retrieves search information by calling the display module;
[0010] The cloud computing platform is used to process information and perform file retrieval: First, it performs initial and secondary screening to conduct the first round of file search and retrieval, and sends a file display signal to the retrieval and display module; then, it analyzes and processes the file information and retrieval information to obtain retrieval accuracy coefficient, update latency coefficient, sorting quality coefficient, and page response coefficient, and then constructs a coefficient matrix to obtain the file retrieval effect evaluation coefficient; then, it sets an abnormal interval and performs interval comparison: when the file retrieval effect is determined to be abnormal, an abnormal signal is generated; when the file retrieval effect is determined to be normal, a management signal is generated;
[0011] The retrieval and display module is used to retrieve and search files and display abnormal alarms: it retrieves and sorts the files in the file retrieval database by receiving file display signals; it performs system optimization and maintenance by receiving abnormal signals; and it performs refined retrieval by receiving management signals.
[0012] Furthermore, the specific process of information processing and file retrieval on the cloud computing platform is as follows:
[0013] Document information includes the frequency of document keywords, the publication time and update time of legal documents, and search information includes the frequency of document clicks, the number of document downloads, and the page response time.
[0014] The frequency of keywords in the document is obtained by extracting and tagging keywords from the document content; the publication and update times of legal documents are obtained by accessing the system backend logs; and search information is collected by system monitoring tools.
[0015] Sa1: First, perform initial screening and extraction using the time code and regional code of legal documents, and then obtain and mark them as initial screening documents;
[0016] Sa2: Re-analyze and process file information:
[0017] Sa2-1: Obtain the content relevance of the initially screened files by the frequency of file keywords and perform secondary screening and extraction, thereby conducting the first round of file search and retrieval, obtaining and marking the files as secondary screened files and sending the file display signal to the retrieval and display module;
[0018] Sa2-2: The search accuracy coefficient is obtained by comparing the keywords entered by the user with the keywords retrieved by the system from the file.
[0019] Sa2-3: Then, by combining the publication time and update time of the legal documents, the system update delay coefficient is obtained;
[0020] Sa3: Then analyze and process the retrieved information:
[0021] Sa3-1: Obtain the sorting quality coefficient by combining file click frequency and file download volume;
[0022] Sa3-2: Obtain the page response coefficient by converting the page response time;
[0023] Sa4: Then, by constructing a coefficient matrix using retrieval accuracy coefficient, update delay coefficient, ranking quality coefficient, and page response coefficient, we can analyze and obtain the evaluation coefficients for document retrieval performance.
[0024] Furthermore, the specific process of initial screening and extraction is as follows:
[0025] Sb1: The default document retrieval database contains N0 legal documents. Any legal document is marked as Fi, the time code of legal document Fi is marked as Tfi, and the region code of legal document Fi is marked as Zfi.
[0026] Sb2: The user sets the time range and geographical range for file retrieval. The time range is marked as T0 and the geographical range is marked as Z0. Then, the time and geographical codes of the legal documents are compared with the time and geographical range conditions of the user's search.
[0027] Sb3: Then, set initial screening criteria to filter legal documents and initially extract N1 legal documents that meet the initial screening criteria:
[0028] Sb3-1: The preset initial screening condition is: if the time code Tfi is within the time range T0 and the area code Zfi is within the area range Z0, then legal document Fi is extracted;
[0029] Sb3-2: Set up a preliminary screening file set, which includes all legal documents that meet the preliminary screening criteria. Mark any file in the preliminary screening file set as a preliminary screening file Cj.
[0030] Furthermore, the specific process of secondary screening and extraction is as follows:
[0031] Users enter keywords to search for files on the search page. The relevance of the N1 initially screened files in the initial file set is calculated based on the keywords, thus performing the first round of file search. The search results are then sorted by content relevance and displayed on the search page. The specific analysis process is as follows:
[0032] Sc1: By scanning the text content of the initial screening file Cj, extracting and marking the file search keywords, and accumulating the total number of all keywords in the text content of the initial screening file Cj, the frequency of the file keywords in the initial screening file Cj is obtained.
[0033] Sc2: Then, calculate the content relevance Dxg of the initially screened file Cj by the frequency of file keywords: Presuppose that there are n0 keywords for file search input by the user. Mark any keyword as G, and mark the cumulative number of keywords G in the initially screened file Cj as the frequency Mg. Obtain the frequency corresponding to the n0 keywords, and then obtain the content relevance Dxg of the initially screened file Cj.
[0034] Sc3: Then obtain the content relevance of N1 initial screening files, and set the rescreening conditions to perform rescreening and extraction, thereby obtaining N2 files that meet the rescreening conditions;
[0035] Sc3-1: The preset rescreening condition is: if the content relevance Dxg of the initially screened file Cj is not equal to 0, then the initially screened file Cj is extracted;
[0036] Sc3-2: Set up a set of files for secondary screening. Include all files that meet the criteria for secondary screening into the set of files for secondary screening. Mark any file in the set of files for secondary screening as a file for secondary screening, Sk.
[0037] Sc3-3: Then, sort the N2 re-screened documents in the re-screened document set in descending order according to their content relevance, and display them in the page search window.
[0038] Furthermore, the specific process of analyzing and processing document information and retrieval information is as follows:
[0039] Sd1: Obtain the search accuracy coefficient:
[0040] The system compares the input keywords with the keywords retrieved from the file. If the input keywords do not match the keywords retrieved from the file, an error signal is generated; otherwise, no signal is generated.
[0041] Then, by comparing keywords in N2 rescreened files, the cumulative amount of detection error signal Hc is obtained, and then the retrieval accuracy coefficient Xzq is obtained;
[0042] Sd2: Get the update latency coefficient:
[0043] The publication time of legal document Fi is marked as Tf, and the update time of legal document Fi is marked as Tg. By combining the publication time Tf and the update time Tg of legal document Fi, the delay time of legal document Fi is analyzed and obtained. Then, by combining the delay times of N0 legal documents, the update delay coefficient Xcy of the system is obtained.
[0044] Sd3: Get the sorting quality coefficient.
[0045] By monitoring access to N2 screened files in the page search window, the corresponding file click frequency and file download volume are obtained;
[0046] The frequency of file clicks on the rescreened file Sk is marked as Wk, the number of file downloads of the rescreened file Sk is marked as Lk, the retrieval quality index Zk of the rescreened file Sk is obtained, and then the ranking quality coefficient Xpx is obtained by combining the retrieval quality indices of N2 rescreened files.
[0047] Sd4: Get the page response time rating.
[0048] Mark the page response time as Tx, assign a corresponding conversion coefficient to the page response time Tx, and obtain the page response coefficient Xxy.
[0049] Furthermore, the specific process of constructing a coefficient matrix to obtain the evaluation coefficients for document retrieval performance is as follows:
[0050] Se1: Preset n1 time nodes, calculate the above four coefficients at the time nodes, and obtain n1 sorting quality coefficients Xpx, page response coefficients Xxy, update latency coefficients Xcy and retrieval accuracy coefficients Xzq respectively, and construct coefficient matrix P;
[0051] Se2: Obtain the eigenvalues of the eigenvectors in the matrix by taking the variance of the column vectors, and label the eigenvalues corresponding to the sorting quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient as q1, q2, q3 and q4 respectively;
[0052] Se3: Then, by calculating the proportion of any feature value in the sum of the four feature values, the weight factors corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient are obtained respectively, and marked as g1, g2, g3 and g4 respectively.
[0053] Se4: Then, by combining the feature values corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient with the weighting factors, the system fluctuation evaluation coefficient BD is obtained.
[0054] Se5: The mean values of n1 sorting quality coefficients, page response coefficients, update latency coefficients, and retrieval accuracy coefficients are obtained, and combined with the system fluctuation evaluation coefficient BD, the file retrieval performance evaluation coefficient Xjs is obtained.
[0055] Furthermore, the specific process for generating abnormal signals and management signals is as follows:
[0056] The abnormal range of the file retrieval effect evaluation coefficient Xjs is set. When the file retrieval effect evaluation coefficient Xjs is in the abnormal range, the file retrieval effect is determined to be abnormal, an abnormal signal is generated and sent to the retrieval and display module, and the system is optimized and maintained. Conversely, the file retrieval effect is determined to be normal, a management signal is generated and sent to the retrieval and display module, and a refined retrieval is performed.
[0057] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0058] This invention combines initial and secondary screening to retrieve and sort legal documents, ensuring their timeliness and regional relevance. Page sorting is based on content relevance calculations, making user retrieval more efficient and convenient. A cloud computing platform analyzes document and retrieval information to assess retrieval accuracy, document update delays, sorting quality, and page responsiveness. This comprehensive analysis evaluates the retrieval effect, enabling monitoring, early warning, and optimization of the retrieval system, ensuring its continuous stability. Furthermore, refined, in-depth retrieval improves the targeting and efficiency of legal document searches. Attached Figure Description
[0059] Figure 1 A schematic diagram of the module of the present invention is shown;
[0060] Figure 2 A schematic diagram of the process of the present invention is shown. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0062] Example 1:
[0063] like Figure 1-2 As shown, the cloud-based legal service platform retrieval system includes a document retrieval database, an information collection module, a cloud computing platform, and a retrieval and display module, wherein the document retrieval database, the cloud computing platform, the information collection module, and the retrieval and display module are connected by signals.
[0064] The work steps are as follows:
[0065] S1: Document retrieval database for storing and updating legal documents: legal documents are retrieved and stored via background upload, time and region codes are set and marked, and the legal documents are dynamically updated;
[0066] By digitally encoding the publication time of legal documents, the time code of the legal documents is obtained. A regional coding database is pre-built, and the applicable regions of the legal documents are coded. For example, if the publication time of a legal document is Tfb, the applicable region of the legal document is φ, and the code of φ in the regional coding database is DQBM, then the time code of the legal document is Tfb and the regional code is DQBM.
[0067] S2: The information acquisition module collects file information and retrieval information: it obtains file information through the file retrieval database and retrieves retrieval information by calling the display module;
[0068] Document information includes the frequency of document keywords, the publication time and update time of legal documents, and search information includes the frequency of document clicks, the number of document downloads, and the page response time.
[0069] The frequency of keywords in the document is obtained by extracting and tagging keywords from the document content; the publication and update times of legal documents are obtained by accessing the system backend logs; and search information is collected by system monitoring tools.
[0070] S3: The cloud computing platform processes information and performs file retrieval: First, it performs initial and secondary screening to conduct the first round of file search and retrieval, and sends a file display signal to the display module; then, it analyzes and processes file information and retrieval information to obtain retrieval accuracy coefficient, update latency coefficient, sorting quality coefficient, and page response coefficient, and then constructs a coefficient matrix to obtain file retrieval effect evaluation coefficients. Next, it sets an abnormal interval and performs interval comparison: when the file retrieval effect is determined to be abnormal, an abnormal signal is generated; when the file retrieval effect is determined to be normal, a management signal is generated. The specific process is as follows:
[0071] Sa1: First, the legal documents are initially screened and extracted using their time and region codes. These documents are then marked as initial screening files. This initial screening ensures the timeliness and regional relevance of the legal documents. The specific process for initial screening is as follows:
[0072] Sb1: The default document retrieval database contains N0 legal documents. Any legal document is marked as Fi, the time code of legal document Fi is marked as Tfi, and the region code of legal document Fi is marked as Zfi.
[0073] Sb2: The user sets the time range and geographical range for file retrieval. The time range is marked as T0 and the geographical range is marked as Z0. Then, the time and geographical codes of the legal documents are compared with the time and geographical range conditions of the user's search.
[0074] Sb3: Then, set initial screening criteria to filter legal documents and initially extract N1 legal documents that meet the initial screening criteria:
[0075] Sb3-1: The preset initial screening condition is: if the time code Tfi is within the time range T0 and the area code Zfi is within the area range Z0, then legal document Fi is extracted;
[0076] Sb3-2: Set up a preliminary screening file set, include all legal documents that meet the preliminary screening criteria into the preliminary screening file set, and mark any file in the preliminary screening file set as a preliminary screening file Cj;
[0077] Sa2: Re-analyze and process file information:
[0078] Sa2-1: Initially screens files based on keyword frequency to determine content relevance and performs secondary screening. Pages are then sorted according to content relevance for more efficient and convenient user retrieval. This initial screening process retrieves and marks files for secondary screening and sends a file display signal to the display module. The specific process of secondary screening and extraction is as follows:
[0079] Users enter keywords to search for files on the search page. The relevance of the N1 initially screened files in the initial file set is calculated based on the keywords, thus performing the first round of file search. The search results are then sorted by content relevance and displayed on the search page. The specific analysis process is as follows:
[0080] Sc1: By scanning the text content of the initial screening file Cj, extracting and marking the file search keywords, and accumulating the total number of all keywords in the text content of the initial screening file Cj, the frequency of the file keywords in the initial screening file Cj is obtained.
[0081] Sc2: Next, calculate the content relevance Dxg of the initially screened file Cj using keyword frequency: Presumably, the user inputs n0 keywords for file searches. Each keyword is labeled G, and the cumulative number of keywords G in the initially screened file Cj is labeled as the frequency Mg. Obtain the frequencies corresponding to the n0 keywords, and then calculate the content relevance Dxg of the initially screened file Cj. The higher the frequency (Mg) of a keyword, the higher the content relevance (Dxg), indicating better content relevance.
[0082] Sc3: Then obtain the content relevance of N1 initial screening files, and set the rescreening conditions to perform rescreening and extraction, thereby obtaining N2 files that meet the rescreening conditions;
[0083] Sc3-1: The preset rescreening condition is: if the content relevance Dxg of the initially screened file Cj is not equal to 0, then the initially screened file Cj is extracted;
[0084] Sc3-2: Set up a set of files for secondary screening. Include all files that meet the criteria for secondary screening into the set of files for secondary screening. Mark any file in the set of files for secondary screening as a file for secondary screening, Sk.
[0085] Sc3-3: Then, sort the N2 re-screened files in the re-screened file set in descending order according to content relevance, and display them in the page search window;
[0086] Sa2-2: The search accuracy coefficient is obtained by comparing the keywords entered by the user with the keywords retrieved by the system from the file.
[0087] Sd1: The specific process for obtaining the retrieval accuracy coefficient is as follows:
[0088] The retrieval accuracy coefficient Xzq is obtained by comparing the input keywords with the keywords in the retrieved file. Keyword comparison is performed using natural language processing. If the input keywords do not match the keywords in the retrieved file, a detection error signal is generated; otherwise, no signal is generated.
[0089] Furthermore, by comparing keywords in N2 re-screened files, the cumulative amount of detection error signals Hc is obtained, and then the retrieval accuracy coefficient Xzq is obtained:
[0090] The higher the cumulative amount Hc of detected error signals, the lower the retrieval accuracy coefficient Xzq, which means a lower retrieval accuracy and a worse file retrieval effect.
[0091] Sa2-3: Then, by combining the publication time and update time of the legal documents, the system update delay coefficient is obtained;
[0092] Sd2: The specific process for obtaining the update delay coefficient is as follows:
[0093] The publication time of legal document Fi is denoted as Tf, and the update time of legal document Fi is denoted as Tg. By combining the publication time Tf and the update time Tg of legal document Fi, the delay time of legal document Fi is analyzed and obtained. Then, by combining the delay times of N0 legal documents, the update delay coefficient Xcy of the system is obtained.
[0094] The publication time of a legal document is the point in time when it is publicly released to the public. The publication time is used to time-encode the legal document, while the update time is the point in time when the system uploads and includes the legal document in the database. The higher the difference between the publication time Tf and the update time Tg, the lower the update delay coefficient Xcy, which indicates that the update time of the legal document is more delayed and the update efficiency is worse, which will result in a worse document retrieval effect.
[0095] Sa3: Then analyze and process the retrieved information:
[0096] Sa3-1: Obtain the sorting quality coefficient by combining file click frequency and file download volume;
[0097] Sd3: The specific process for obtaining the sorting quality coefficient is as follows:
[0098] By monitoring access to N2 screened files in the page search window, the corresponding file click frequency and file download volume are obtained;
[0099] The click frequency of the rescreened file Sk is denoted as Wk, the download volume of the rescreened file Sk is denoted as Lk, the retrieval quality index Zk of the rescreened file Sk is obtained, and then the ranking quality coefficient Xpx is obtained:
[0100]
[0101] Wherein, μ1 and μ2 are the weight factor coefficients of file click frequency Wk and file download volume Lk, respectively, and both μ1 and μ2 are greater than 0; when the file click frequency Wk and file download volume Lk are higher, the ranking quality coefficient Xpx is higher, which means that the retrieval ranking quality is better, which will make the file retrieval effect better;
[0102] Sa3-2: Obtain the page response coefficient by converting the page response time;
[0103] Sd4: The specific process for obtaining the page response coefficient is as follows:
[0104] The page response time is converted to obtain the page response coefficient;
[0105] Mark the page response time as Tx, assign a corresponding conversion coefficient to the page response time Tx, and obtain the page response coefficient Xxy:
[0106] Where ε is the conversion coefficient, and ε is greater than 0. The conversion coefficient is obtained by preset. When the page response time Tx is higher, the page response coefficient Xxy is lower, which means that the page response effect is worse and the file search effect will be worse.
[0107] Sa4: Then, by constructing a coefficient matrix using retrieval accuracy coefficient, update delay coefficient, ranking quality coefficient, and page response coefficient, we can analyze and obtain the evaluation coefficients for file retrieval performance.
[0108] The specific process of constructing a coefficient matrix and analyzing it to obtain the evaluation coefficients for document retrieval performance is as follows:
[0109] Preset n1 time nodes, and calculate the above four coefficients at each time node to obtain n1 sorting quality coefficients Xpx, page response coefficients Xxy, update latency coefficients Xcy and retrieval accuracy coefficients Xzq respectively;
[0110] Construct the coefficient matrix P:
[0111]
[0112] The eigenvalues of the eigenvectors in the matrix are obtained by taking the variance of the column vectors. The eigenvalues corresponding to the sorting quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient are labeled as q1, q2, q3 and q4 respectively.
[0113] The specific calculation formulas for q1, q2, q3, and q4 are as follows:
[0114] It is the average of the n1 sorting quality coefficients;
[0115] This is the average of the response coefficients of n1 pages;
[0116] It is the average of n1 update delay coefficients;
[0117] It is the average of the n1 retrieval accuracy coefficients;
[0118] Then, by calculating the proportion of any feature value in the sum of the four feature values, the weight factors corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient are obtained respectively, and labeled as g1, g2, g3 and g4 respectively.
[0119] For example, g1: Similarly, the specific calculation formulas for g2, g3, and g4 can be derived.
[0120] Then, by combining the feature values corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient, and retrieval accuracy coefficient with weighting factors, the system fluctuation evaluation coefficient BD is obtained:
[0121]
[0122] The higher the eigenvalue corresponding to the coefficient, the more severe the fluctuation of the coefficient at n1 time points, and the more unstable the coefficient. The higher the system fluctuation evaluation coefficient BD, the worse the stability of the combined effect of the four coefficients, which will result in a worse overall stability of the system retrieval and thus a worse file retrieval effect.
[0123] The mean values of n1 sorting quality coefficients, page response coefficients, update latency coefficients, and retrieval accuracy coefficients are calculated and combined with the system fluctuation evaluation coefficient BD to obtain the overall file retrieval performance evaluation coefficient Xjs.
[0124] When the average values of the sorting quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient are higher, and the system fluctuation evaluation coefficient BD is lower, the file retrieval performance evaluation coefficient Xjs is higher, indicating that the file retrieval performance is better.
[0125] The specific process of generating abnormal signals and management signals is as follows:
[0126] The abnormal range of the file retrieval effect evaluation coefficient Xjs is set. When the file retrieval effect evaluation coefficient Xjs is in the abnormal range, the file retrieval effect is determined to be abnormal, and an abnormal signal is generated and sent to the retrieval and display module to optimize and maintain the system and ensure the continuous stability of the system. Conversely, the file retrieval effect is determined to be normal, and a management signal is generated and sent to the retrieval and display module for refined retrieval.
[0127] S4: Retrieve and display the file and display the abnormal alarm: retrieve and sort the files in the file retrieval database by receiving the file display signal; optimize and maintain the system by receiving the abnormal signal; and perform fine-grained retrieval by receiving the management signal.
[0128] The specific process for optimizing and maintaining the system is as follows: the display module receives abnormal signals, immediately edits and displays the alarm text, then arranges for professional personnel to optimize and improve the software algorithm for file retrieval, replaces and upgrades the system hardware, and maintains the system network and transmission.
[0129] The specific process of refined retrieval is as follows:
[0130] Sort the N2 rescreened files in descending order according to their retrieval quality index, extract the top n2 rescreened files by retrieval quality index, set up a valid file set, include all the top n2 rescreened files in the valid file set, and mark any file in the valid file set as a valid file Yu.
[0131] Identify the text content of the valid document Yu and extract the high-frequency words. The extraction process is as follows: mark any word as a candidate word, extract the candidate word from all text content, accumulate the frequency of the candidate words, sort them in descending order of frequency, and mark the candidate words in the top n3 positions as high-frequency words.
[0132] The n3 high-frequency words are integrated as subject terms, and a deep search is performed in the form of "keyword + subject term". The retrieval and display module sends a call command to the document retrieval database. The information collection module and the cloud computing platform sequentially re-collect and recalculate the content relevance of legal documents in the document retrieval database, thereby performing a second round of document search. The search page is displayed based on the recalculated content relevance. The deep search improves the retrieval efficiency of legal documents.
[0133] The process of refined retrieval is similar to that of the initial search, but the difference lies in expanding the content of the original "keywords" and increasing the restrictions on the search process for legal documents. This narrows the range of extractable legal documents and reduces the number of legal documents retrieved, thus achieving refined retrieval. Furthermore, it analyzes the user's viewing behavior on the page files after the initial search, making the search more targeted and efficient.
[0134] In summary, this invention combines initial screening and secondary screening to retrieve and sort legal documents, ensuring their timeliness and regional relevance. Page sorting based on content relevance makes retrieval more efficient and convenient for users. Cloud computing platforms analyze document and retrieval information to assess retrieval accuracy, document update delays, sorting quality, and page responsiveness. A comprehensive analysis of the retrieval results enables monitoring, early warning, and optimization of the retrieval system, ensuring its continuous stability. Furthermore, refined, in-depth retrieval improves the targeting and efficiency of legal document searches.
[0135] The size of the interval and threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by those skilled in the art for each set of sample data; as long as it does not affect the ratio between the parameter and the quantized value.
[0136] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0137] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A cloud-based legal service platform retrieval system, characterized in that: It includes a file retrieval database, an information acquisition module, a cloud computing platform, and a retrieval and display module, wherein the file retrieval database, cloud computing platform, information acquisition module, and retrieval and display module are connected by signals; The document retrieval database is used to store and update legal documents: legal documents are retrieved and stored through background uploads, time and region codes are set and marked, and legal documents are dynamically updated; The information acquisition module is used to collect file information and retrieve search information: it obtains file information through a file retrieval database and retrieves search information by calling the display module; The cloud computing platform is used to process information and perform file retrieval: First, it performs initial and secondary screening to conduct the first round of file search and retrieval, and sends a file display signal to the retrieval and display module; then, it analyzes and processes the file information and retrieval information to obtain retrieval accuracy coefficient, update latency coefficient, sorting quality coefficient, and page response coefficient, and then constructs a coefficient matrix to obtain the file retrieval effect evaluation coefficient; then, it sets an abnormal interval and performs interval comparison: when the file retrieval effect is determined to be abnormal, an abnormal signal is generated and sent to the retrieval and display module; when the file retrieval effect is determined to be normal, a management signal is generated and sent to the retrieval and display module. The retrieval and display module is used to retrieve and search files and display abnormal alarms: it retrieves and sorts the files in the file retrieval database by receiving file display signals; it performs system optimization and maintenance by receiving abnormal signals; and it performs refined retrieval by receiving management signals. Document information includes the frequency of document keywords, the publication and update times of legal documents, and search information includes the frequency of document clicks, the number of document downloads, and the page response time.
2. The cloud-based legal service platform retrieval system according to claim 1, characterized in that: The specific process of information processing and file retrieval on a cloud computing platform is as follows: The frequency of keywords in the document is obtained by extracting and tagging keywords from the document content; the publication and update times of legal documents are obtained by accessing the system backend logs; and search information is collected by system monitoring tools. Sa1: First, perform initial screening and extraction using the time code and regional code of legal documents, and then obtain and mark them as initial screening documents; Sa2: Re-analyze and process file information: Sa2-1: Obtain the content relevance of the initially screened files by the frequency of file keywords and perform secondary screening and extraction, thereby conducting the first round of file search and retrieval, obtaining and marking the files as secondary screened files and sending the file display signal to the retrieval and display module; Sa2-2: The search accuracy coefficient is obtained by comparing the keywords entered by the user with the keywords retrieved by the system from the file. Sa2-3: Then, by combining the publication time and update time of the legal documents, the system update delay coefficient is obtained; Sa3: Then analyze and process the retrieved information: Sa3-1: Obtain the sorting quality coefficient by combining file click frequency and file download volume; Sa3-2: Obtain the page response coefficient by converting the page response time; Sa4: Then, by constructing a coefficient matrix using retrieval accuracy coefficient, update delay coefficient, ranking quality coefficient, and page response coefficient, we can analyze and obtain the evaluation coefficients for document retrieval performance.
3. The cloud-based legal service platform retrieval system according to claim 2, characterized in that: The specific process of initial screening and extraction is as follows: Sb1: The default document retrieval database contains N0 legal documents. Any legal document is marked as Fi, the time code of legal document Fi is marked as Tfi, and the region code of legal document Fi is marked as Zfi. Sb2: The user sets the time range and geographical range for file retrieval. The time range is marked as T0 and the geographical range is marked as Z0. Then, the time and geographical codes of the legal documents are compared with the time and geographical range conditions of the user's search. Sb3: Then, set initial screening criteria to filter legal documents and initially extract N1 legal documents that meet the initial screening criteria: Sb3-1: The preset initial screening condition is: if the time code Tfi is within the time range T0 and the area code Zfi is within the area range Z0, then legal document Fi is extracted; Sb3-2: Set up a preliminary screening file set, which includes all legal documents that meet the preliminary screening criteria. Mark any file in the preliminary screening file set as a preliminary screening file Cj.
4. The cloud-based legal service platform retrieval system according to claim 3, characterized in that: The specific process of secondary screening and extraction is as follows: Users enter keywords to search for files on the search page. The relevance of the N1 initially screened files in the initial file set is calculated based on the keywords, thus performing the first round of file search. The search results are then sorted by content relevance and displayed on the search page. The specific analysis process is as follows: Sc1: By scanning the text content of the initial screening file Cj, extracting and marking the file search keywords, and accumulating the total number of all keywords in the text content of the initial screening file Cj, the frequency of the file keywords in the initial screening file Cj is obtained. Sc2: Then, calculate the content relevance Dxg of the initially screened file Cj by the frequency of file keywords: Presuppose that there are n0 keywords for file search input by the user. Mark any keyword as G, and mark the cumulative number of keywords G in the initially screened file Cj as the frequency Mg. Obtain the frequency corresponding to the n0 keywords, and then obtain the content relevance Dxg of the initially screened file Cj. Sc3: Then obtain the content relevance of N1 initial screening files, and set the rescreening conditions to perform rescreening and extraction, thereby obtaining N2 files that meet the rescreening conditions; Sc3-1: The preset rescreening condition is: if the content relevance Dxg of the initially screened file Cj is not equal to 0, then the initially screened file Cj is extracted; Sc3-2: Set up a set of files for secondary screening. Include all files that meet the criteria for secondary screening into the set of files for secondary screening. Mark any file in the set of files for secondary screening as a file for secondary screening, Sk. Sc3-3: Then, sort the N2 rescreened files in the rescreened file set in descending order according to their content relevance, and display them in the page search window.
5. The cloud-based legal service platform retrieval system according to claim 4, characterized in that: The specific process of analyzing and processing document information and retrieval information is as follows: Sd1: Obtain the search accuracy coefficient: The search is performed by comparing the input keywords with the keywords in the retrieved files. If the input keywords do not match the keywords in the retrieved files, a detection error signal is generated; otherwise, no signal is generated. Then, by comparing the keywords of N2 screened files, the cumulative amount of detection error signals Hc is obtained, and the search accuracy coefficient Xzq is obtained. Sd2: Get the update latency coefficient: The publication time of legal document Fi is marked as Tf, and the update time of legal document Fi is marked as Tg. By combining the publication time Tf and the update time Tg of legal document Fi, the delay time of legal document Fi is analyzed and obtained. Then, by combining the delay times of N0 legal documents, the update delay coefficient Xcy of the system is obtained. Sd3: Get the sorting quality coefficient. By monitoring access to N2 screened files in the page search window, the corresponding file click frequency and file download volume are obtained; the file click frequency of screened file Sk is marked as Wk, and the file download volume of screened file Sk is marked as Lk, and the search quality index Zk of screened file Sk is obtained. Then, by combining the search quality indices of N2 screened files, the ranking quality coefficient Xpx is obtained. Sd4: Get the page response time rating. Mark the page response time as Tx, assign a corresponding conversion coefficient to the page response time Tx, and obtain the page response coefficient Xxy.
6. The cloud-based legal service platform retrieval system according to claim 5, characterized in that: The specific process of constructing a coefficient matrix and analyzing it to obtain the evaluation coefficients for document retrieval performance is as follows: Se1: Preset n1 time nodes, calculate the above four coefficients at the time nodes, and obtain n1 sorting quality coefficients Xpx, page response coefficients Xxy, update latency coefficients Xcy and retrieval accuracy coefficients Xzq respectively, and construct coefficient matrix P; Se2: Obtain the eigenvalues of the eigenvectors in the matrix by taking the variance of the column vectors, and label the eigenvalues corresponding to the sorting quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient as q1, q2, q3 and q4 respectively; Se3: Then, by calculating the proportion of any feature value in the sum of the four feature values, the weight factors corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient are obtained respectively, and marked as g1, g2, g3 and g4 respectively. Se4: Then, by combining the feature values corresponding to the ranking quality coefficient, page response coefficient, update latency coefficient and retrieval accuracy coefficient with the weighting factors, the system fluctuation evaluation coefficient BD is obtained. Se5: The mean values of n1 sorting quality coefficients, page response coefficients, update latency coefficients, and retrieval accuracy coefficients are obtained, and combined with the system fluctuation evaluation coefficient BD, the file retrieval performance evaluation coefficient Xjs is obtained.
7. The cloud-based legal service platform retrieval system according to claim 6, characterized in that: The specific process of generating abnormal signals and management signals is as follows: The abnormal range of the file retrieval effect evaluation coefficient Xjs is set. When the file retrieval effect evaluation coefficient Xjs is in the abnormal range, the file retrieval effect is determined to be abnormal, an abnormal signal is generated and sent to the retrieval and display module, and the system is optimized and maintained. Conversely, the file retrieval effect is determined to be normal, a management signal is generated and sent to the retrieval and display module, and a refined retrieval is performed.
Citation Information
Patent Citations
Method and system for retrieval of judicial cases
CN107247743A
Local archive search management system and method based on digital humanity
CN116594957A