Real estate registration archive arrangement service method and system

By dividing high-frequency and low-frequency queries of real estate registration files, filtering high-frequency stable query users and calculating query efficiency, and identifying appropriate classification tags, the problem of low query efficiency caused by unreasonable archive tags is solved, and the user query efficiency is improved.

CN120386901AActive Publication Date: 2025-07-29YUNNAN GAOYANG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510887203.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

In the process of organizing real estate registration files in the existing technology, there is a lack of effective methods to screen users with high-frequency and stable queries, resulting in unreasonable classification of file tags and inefficient user query.

Method used

By counting the relative query ratio of real estate registration files, it is divided into high-frequency and low-frequency query files, users with high-frequency stable query are selected, query efficiency is calculated and divided into efficient and inefficient users, the differences between user query information and archive classification labels are analyzed, and keywords that meet multiple user types are identified as new classification labels.

Benefits of technology

The efficiency of real estate registration file query has been improved, and the speed and accuracy of user query has been improved through reasonable file tag classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386901A_ABST
    Figure CN120386901A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of archive arrangement, and provides a real estate registration archive arrangement service method and system, and the method comprises the steps: dividing a real estate registration archive into a high-frequency query archive and a low-frequency query archive, and screening out a high-frequency stable query user corresponding to the high-frequency query archive, by calculating the query efficiency of high-frequency query archives, the high-frequency query archives are divided into high-efficiency archives and low-efficiency archives, and users corresponding to the high-efficiency archives are divided into high-efficiency users and low-efficiency users; whether the classification label of the archive is changed or not is judged by analyzing the difference between the query information of the efficient users and the low-efficiency users and the archive and the number of the efficient users, and keywords conforming to the efficient users and the low-efficiency users are identified as new classification labels. Therefore, the problem of low user query efficiency caused by unreasonable tag classification of the archives in the real estate registration archive arrangement process is solved, and the user query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of file arrangement, and specifically relates to a method and system for arranging real estate registration files. Background Art

[0002] There are many problems in the existing technology in terms of grasping the core user needs and file tag classification. For example, for users who frequently query files, there is a lack of effective methods to screen out users with high-frequency and stable queries, making it impossible to deeply understand the needs of these core users, and it is difficult to provide targeted services and optimize file management strategies, resulting in inaccurate file tag classification. In terms of file tag classification, the existing technology may lack scientific and effective methods to accurately classify file tags, making the classified tags of files not match the actual query needs of users, which may lead to long query time, low accuracy, and low query efficiency for users.

[0003] Therefore, the present invention provides a method and system for arranging real estate registration files. Summary of the Invention

[0004] In order to make up for the deficiencies of the existing technology and solve at least one of the technical problems proposed in the background art.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows: Step 1: Count the number of queries of each real estate registration file within the historical period and the total number of real estate files, calculate the relative query ratio of the real estate registration files, and divide the real estate registration files into high-frequency query files and low-frequency query files; Step 2: For high-frequency query files, screen out users with high-frequency and stable queries and establish a corresponding user library; Step 3: Calculate the query efficiency of high-frequency query files, compare it with a threshold, mark the files with query efficiency greater than or equal to the threshold as efficient files, and divide the corresponding users in the user library into efficient users and inefficient users for the efficient files; Step 4: Judge whether the difference between the query information of efficient users and inefficient users and the classified tags of the files is large. If it is large, judge whether it is necessary to change the classified tags according to the ratio of the actual number of efficient users to the ideal number of efficient users; Step 5: If the classified tags need to be changed, establish a query information library based on efficient users and inefficient users, extract keywords, and identify the keywords that meet both efficient users and inefficient users as the new classified tags.

[0006] Further, the method of dividing the real estate registration files into high-frequency query files and low-frequency query files is as follows: Compare the relative query ratio of the real estate registration files with a threshold; Classify the real estate registration files with a relative query ratio greater than or equal to the threshold as high-frequency query files; Classify the real estate registration files with a relative query ratio less than the threshold as low-frequency query files.

[0007] Furthermore, the calculation method of the relative query ratio of the real estate registration files is as follows: Within the historical period, calculate the ratio of the number of queries of the real estate registration files to the total number of real estate registration files to obtain the relative query ratio of the real estate registration files.

[0008] Furthermore, the process of screening out users with high-frequency and stable queries includes: Divide the historical period into n small time segments. For each user, count the number of queries in each small time segment to generate a query count sequence; Compare the number of queries of each user in the small time segment with the set threshold. If the number of queries is greater than the threshold in a small time segment, mark this small time segment as a high-frequency time segment; Count the number of high-frequency time segments of each user in the historical period, and calculate the ratio with the total number of small time segments to obtain the high-frequency time segment ratio; For each user, compare the high-frequency time segment ratio and the standard deviation of the query count sequence with the threshold respectively, and screen out the users with a high-frequency time segment ratio greater than or equal to the threshold and a query time standard deviation greater than or equal to the threshold as high-frequency and stable query users.

[0009] Furthermore, the calculation process of the query efficiency of the high-frequency query files includes: Record the start time and end time of each query of the real estate registration files, calculate the time consumption of each query, then add up the time consumptions of all query operations within a certain period of time, and divide by the number of queries to obtain the average query time; Count the ratio of the number of times of successfully querying the required files to the total number of queries within a certain time to obtain the query success rate; A successful query means that the user can accurately obtain the required file information within the specified time; When the user inputs keywords or other retrieval conditions for query, calculate the ratio of the number of files related to the user's actual needs in the retrieval results to the total number of retrieval results to obtain the retrieval result accuracy rate; Standardize the average query time, query success rate, and retrieval result accuracy rate. Multiply the standardized data by their respective weights, and then add them up to obtain the query efficiency.

[0010] Furthermore, the process of classifying the corresponding users in the user library into efficient users and inefficient users includes: For each user corresponding to efficient files, count the number of efficient files among the files queried by the user within the historical period and the total number of queries, and calculate the ratio of efficient files; Compare the ratio of efficient files of each user with the set threshold; If the ratio of efficient files is greater than or equal to the threshold, the user is an efficient user; If the ratio of efficient files is less than the threshold, the user is an inefficient user.

[0011] Furthermore, the method for determining whether the difference between the query information of efficient users and inefficient users and the classification labels of files is large is as follows: Collect the query information of efficient users and inefficient users, extract keywords, and generate query information keyword sets A and classification label keyword sets B respectively; Calculate the intersection and union of set A and set B respectively, and count the number of elements in the intersection and union; Calculate the matching degree score between the query information of efficient users and inefficient users and the file classification labels according to the formula of Jaccard coefficient; Compare the matching degree score between the query information of efficient users and inefficient users and the file classification labels with the threshold; If the matching degree score is less than or equal to the threshold, it means that the difference between the query information of efficient users and inefficient users and the file classification labels is large.

[0012] Furthermore, the method for identifying keywords that conform to both efficient users and inefficient users is as follows: Label the extracted keywords to clarify whether each keyword belongs to an efficient user or an inefficient user, and organize the extracted keywords and the corresponding user types into the form of a data set; Use the Naive Bayes model to identify keywords that conform to both efficient users and inefficient users.

[0013] Furthermore, the process of using the Naive Bayes model to identify keywords that conform to both efficient users and inefficient users includes: Calculate the prior probability and conditional probability of efficient users and inefficient users respectively; Compare the product of the prior probability and conditional probability of efficient users with the product of the prior probability and conditional probability of inefficient users; If the product of the prior probability and conditional probability of efficient users is approximately equal to the product of the prior probability and conditional probability of inefficient users, it means that the keyword conforms to both efficient users and inefficient users.

[0014] A real estate registration file sorting service system includes the following modules: File query frequency classification module: Count the number of queries for each real estate registration file within the historical period and the total number of real estate files, calculate the relative query ratio of the real estate registration files, and classify the real estate registration files into high-frequency query files and low-frequency query files; High-frequency stable user screening module: For high-frequency query files, screen out users with high-frequency stable queries and establish a corresponding user database; Efficient file and user classification module: Calculate the query efficiency of high-frequency query files, compare it with the threshold, mark the files with query efficiency greater than or equal to the threshold as efficient files, and for efficient files, classify the corresponding users in the user database into efficient users and inefficient users; Classification label difference evaluation module: Judge whether the difference between the query information of efficient users and inefficient users and the classification labels of the files is large. If it is large, judge whether the classification labels need to be changed according to the ratio of the actual number of efficient users to the ideal number of efficient users; New classification label generation module: If the classification labels need to be changed, establish a query information database based on efficient users and inefficient users, extract keywords, and identify the keywords that meet both efficient users and inefficient users as the new classification labels.

[0015] The beneficial effects of the present invention are as follows: By classifying real estate registration files into high-frequency query files and low-frequency query files, screening out high-frequency stable query users corresponding to high-frequency query files, calculating the query efficiency of high-frequency query files, classifying high-frequency query files into efficient files and inefficient files, classifying the users corresponding to efficient files into efficient users and inefficient users, analyzing the difference between the query information of efficient users and inefficient users and the files, and the number of efficient users to judge whether to change the classification labels of the files, and identifying the keywords that meet both efficient users and inefficient users as the new classification labels, the problem that the label classification of files is unreasonable during the process of sorting real estate registration files, resulting in low efficiency when users query, is solved, and the query efficiency of users is improved. Brief description of the drawings

[0016] The present invention will be further described below with reference to the drawings.

[0017] Figure 1 is the step flowchart of a method for sorting real estate registration file services according to an embodiment of the present invention; Figure 2 is the program block diagram of a system for sorting real estate registration file services according to an embodiment of the present invention. Detailed implementation manners

[0018] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation manners.

[0019] Example 1 Please refer to Figure 1 As shown, a method for sorting real estate registration files according to the embodiments of the present invention includes the following steps: Step 1: Count the number of inquiries of each real estate registration file within the historical period and the total number of real estate files, calculate the relative inquiry ratio of the real estate registration files, and divide the real estate registration files into high-frequency inquiry files and low-frequency inquiry files; In Step 1, the relative inquiry ratio refers to the ratio of the number of inquiries of real estate registration files to the total number of real estate registration files within the historical period. The calculation formula is:

[0020] In Step 1, the process of dividing the real estate registration files into high-frequency inquiry files and low-frequency inquiry files includes: Compare the relative inquiry ratio of the real estate registration files with the threshold value. Divide the real estate registration files with a relative inquiry ratio greater than or equal to the threshold value into high-frequency inquiry files, and divide the real estate registration files with a relative inquiry ratio less than the threshold value into low-frequency inquiry files; Among them, the setting of the threshold value of the relative inquiry ratio includes: calculating the relative inquiry ratio of all real estate registration files within the historical period, and calculating the average value. Use the average value of the relative inquiry ratio as the threshold value; Exemplarily, assume that File A has been inquired 30 times in a year. According to the relative inquiry ratio calculation formula, the relative inquiry ratio of File A is: 30÷1000 = 0.03. File B has been inquired 10 times in a year, and its relative inquiry ratio is 10÷1000 = 0.01. File C has been inquired 50 times in a year, and its relative inquiry ratio is 50÷1000 = 0.05, and so on. Count the number of inquiries of 1000 files and calculate their relative inquiry ratios respectively; Assume that after calculating the relative inquiry ratios of 1000 files, add up all the relative inquiry ratios and divide by 1000 to get an average value of 0.02. Use 0.02 as the threshold value of the relative inquiry ratio; For File A, its relative inquiry ratio is 0.03, 0.03>0.02, so File A is divided into high-frequency inquiry files; For File B, its relative inquiry ratio is 0.01, 0.01<0.02, so File B is divided into low-frequency inquiry files; For File C, its relative inquiry ratio is 0.05, 0.05>0.02, so File C is divided into high-frequency inquiry files; Step 2: For high-frequency query archives, screen out users with high-frequency and stable queries and establish a corresponding user library; In Step 2, the process of screening out users with high-frequency and stable queries is as follows: Divide the historical period into n small time segments. For each user, count the number of queries in each small time segment to generate a query count sequence; Compare the number of queries of each user in the small time segment with the set threshold. If the number of queries is greater than the threshold in a small time segment, mark this small time segment as a high-frequency time segment; Among them, the setting of the query count threshold includes: calculating the average value of the query count sequence and using the average value as the query count threshold; Count the number of high-frequency time segments of each user in the historical period, calculate the ratio with the total number of small time segments to obtain the high-frequency time segment ratio; Calculate the standard deviation of the query count sequence of each user. The formula is: , where n is the number of small time segments, is the number of queries in each small time segment, is the average value of the query count; For each user, compare the high-frequency time segment ratio and the standard deviation of the query count sequence with the threshold respectively. Screen out users with a high-frequency time segment ratio greater than or equal to the threshold and a query time standard deviation greater than or equal to the threshold as high-frequency and stable query users; Step 3: Calculate the query efficiency of high-frequency query archives, compare it with the threshold, mark the archives with query efficiency greater than or equal to the threshold as efficient archives. For efficient archives, divide the corresponding users in the user library into efficient users and inefficient users; In Step 3, the calculation process of the query efficiency of high-frequency query archives includes: Record the start time and end time of each query of the real estate registration archives, calculate the time consumption of each query, then add up the time consumptions of all query operations within a certain period (such as one month, one quarter), and divide by the number of queries to obtain the average query time; Count the ratio of the number of times of successfully querying the required archives to the total number of queries within a certain time to obtain the query success rate. Among them, a successful query means that the user can accurately obtain the required archive information within the specified time; When the user inputs keywords or other retrieval conditions for query, calculate the ratio of the number of archives related to the user's actual needs in the retrieval results to the total number of retrieval results to obtain the retrieval result accuracy rate; Comprehensively calculate the query efficiency according to the average query time, query success rate and retrieval result accuracy rate. The process includes: Normalize the data of the average query time, query success rate, and retrieval result accuracy. For the average query time, the smaller the value, the better. Use the formula to normalize it and obtain the normalized value of the average query time, where X is the original average query time, Max is the maximum value of the average query time, and Min is the minimum value of the average query time; For the query success rate and retrieval result accuracy, the larger the value, the better. Use the formula to normalize them and obtain the normalized values of the query success rate and retrieval result accuracy, where X is the original query success rate or retrieval result accuracy, Max is the maximum value of the query success rate or retrieval result accuracy, and Min is the minimum value of the query success rate or retrieval result accuracy; Multiply the normalized data by their respective weights and then add them up to obtain the comprehensive query efficiency. The formula is: Among them, 、 、 are the weights of the normalized value of the average query time, the normalized value of the query success rate, and the normalized value of the retrieval result accuracy respectively; Compare the calculated query efficiency of each file with the set threshold. If the query efficiency of the file is higher than the threshold, classify it as an efficient file. If the query efficiency of the file is lower than the threshold, classify it as an inefficient file; Among them, the setting of the threshold for file query efficiency includes: adding up the query efficiencies of all files to obtain the total query efficiency, and dividing the total query efficiency by the number of files to obtain the average value of file query efficiency. The calculation formula is: , and use the average value of the query efficiency as the threshold, where is the average value, is the query efficiency of the i-th file, and n is the total number of files; In step three, the process of classifying the users corresponding to the high-efficiency query files into high-efficiency users and low-efficiency users includes: For each user corresponding to an efficient file, count the number N 高效 of files that belong to efficient files and the total number of queries N 总 in the files queried by this user during the historical period, and calculate the ratio of efficient files occupied: , where R i is the ratio of efficient files occupied; Compare the ratio of efficient files occupied by each user with the set threshold. If the ratio of efficient files occupied is greater than or equal to the threshold, the user is a high-efficiency user. If the ratio of efficient files occupied is less than the threshold, the user is a low-efficiency user; Among them, the threshold setting of the high-efficiency file occupancy ratio includes: listing the high-efficiency file occupancy ratio distribution of all users, and taking the median of the high-efficiency file occupancy ratio as the threshold of the high-efficiency occupancy ratio; Exemplarily, assume that there are currently 5 frequently queried files, numbered A, B, C, D, and E, and there are three users, numbered 1, 2, and 3. The data parameters of querying these files in the past quarter (historical period) include: | File | Total Query Times | Query Success Times | Total Time Consumed | Number of Retrieval Results | Number of Relevant Retrieval Results |; | A | 10 times | 8 times | 30 minutes | 20 pieces | 16 pieces |; | B | 8 times | 7 times | 24 minutes | 15 pieces | 12 pieces |; | C | 12 times | 10 times | 48 minutes | 25 pieces | 20 pieces |; | D | 6 times | 5 times | 18 minutes | 12 pieces | 9 pieces |; | E | 9 times | 8 times | 27 minutes | 18 pieces | 15 pieces |; Calculate the average query time: File A: The average query time is 30÷10 = 3 minutes; File B: The average query time is 24÷8 = 3 minutes; File C: The average query time is 48÷12 = minutes; File D: The average query time is 18÷6 = 3 minutes; File E: The average query time is 27÷9 = 3 minutes; Calculate the query success rate: File A: The query success rate is 8÷10 = 0.8; File B: The query success rate is 7÷8 = 0.875; File C: The query success rate is 10÷12≈0.833; File D: The query success rate is 5÷6≈0.833; File E: The query success rate is 8÷9≈0.889; Calculate the retrieval result accuracy rate: File A: The retrieval result accuracy rate is 16÷20 = 0.8; File B: The retrieval result accuracy rate is 12÷15 = 0.8; File C: The retrieval result accuracy rate is 20÷25 = 0.8; File D: The retrieval result accuracy rate is 9÷12 = 0.75; File E: The retrieval result accuracy rate is 15÷18≈0.833; Assume that the maximum value of the average query time is Max = 4 minutes, the minimum value is Min = 3 minutes, the maximum value of the query success rate is Max = 0.889, the minimum value is Min = 0.8, and the maximum value of the retrieval result accuracy rate is Max = 0.833, the minimum value is Min = 0.75. Let the weights , , ; After calculation, the comprehensive query efficiencies of File A, File B, File C, File D, and File E are 0.58, 0.712, 0.291, 0.511, and 1 respectively. The total query efficiency = 0.58 + 0.712 + 0.291 + 0.511 = 3.094. The average query efficiency is 3.094 ÷ 5 = 0.6188, which is used as the threshold; Files B (0.712) and E (1) with efficiencies higher than the threshold are high-efficiency files, and Files A (0.58), C (0.291), and D (0.511) are low-efficiency files; Statistical user query situation: User 1: The total number of queries N 总1 = 15 times, among which the number of queries for high-efficiency files (Files B and E) is N 高效1 = 8 times, and the proportion of high-efficiency files ; User 2: The total number of queries N 总2 = 12 times, among which the number of queries for high-efficiency files (Files B and E) is N 高效2 = 4 times, and the proportion of high-efficiency files ; User 3: The total number of queries N 总3 = 18 times, among which the number of queries for high-efficiency files (Files B and E) is N 高效3 = 10 times, and the proportion of high-efficiency files ; Set the threshold for the proportion of high-efficiency files and classify users: The proportions of high-efficiency files for all users are 0.533, 0.333, and 0.556 respectively. Sorted from small to large, they are 0.333, 0.533, and 0.556. The median is 0.533, which is used as the threshold for the proportion of high-efficiency files; The proportion of high-efficiency files of User 1 is equal to the threshold, so User 1 is a high-efficiency user. The proportion of high-efficiency files of User 2 is less than the threshold, so User 2 is a low-efficiency user. The proportion of high-efficiency files of User 3 is greater than the threshold, so User 3 is a high-efficiency user; Step 4: Determine whether the difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the files is large. If it is large, judge whether the classification labels need to be changed according to the ratio of the actual number of high-efficiency users to the ideal number of high-efficiency users; In Step 4, the process of determining whether the difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the files is large includes: Collect the query information of high-efficiency users and low-efficiency users, where the query information includes keywords, phrases or sentences entered by users during queries, and obtain the classification labels of the archives. Use the lexical analysis tool in natural language processing technology to process the query information of high-efficiency users and low-efficiency users as well as the classification labels of the archives, extract keywords, and respectively generate the query information keyword set A and the classification label keyword set B. Calculate the intersection of set A and set B, that is, the set composed of elements that belong to both A and B, denoted as Count the number of elements in the intersection, denoted as ; Calculate the union of set A and set B, that is, the set composed of elements that belong to A or belong to B, denoted as Count the number of elements in the union, denoted as ; According to the formula of Jaccard similarity coefficient: Calculate the matching degree score between the query information of high-efficiency users and low-efficiency users and the archive classification labels. Among them, the Jaccard similarity coefficient is an index used to measure the similarity between two sets, and its value range is [0, 1]. The closer the coefficient is to 1, the higher the similarity between the two sets, and the closer it is to 0, the lower the similarity. Compare the matching degree score between the query information of high-efficiency users and low-efficiency users and the archive classification labels with the threshold. If the matching degree score is less than or equal to the threshold, it means that there is a large difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the archives. Exemplarily, assume there are 3 high-efficiency users (User A, User B, User C), 4 low-efficiency users (User D, User E, User F, User G), and there are 8 real estate registration archives in the system, namely Archive 1, Archive 2, Archive 3, Archive 4, Archive 5, Archive 6, Archive 7, Archive 8. The query information and archive classification labels are as follows: Query information of high-efficiency users: User A: "Query the latest archives of residential property registration in Haidian District" User B: "Find the archives of commercial real estate mortgage registration in Chaoyang District" User C: "Search for the archives of industrial land change registration after 2020" Query information of low-efficiency users: User D: "I want to query the archives related to houses" User E: "Find the archives related to real estate registration" User F: "Are there any archives related to mortgages?" User G: "Query the previous registration archives" File Classification Tags: File 1: "Residential Property Rights Registration in Haidian District - Handled in 2018" File 2: "Commercial Real Estate Mortgage Registration in Chaoyang District - Handled in 2016" File 3: "Industrial Land Alteration Registration in Xicheng District - Handled in 2015" File 4: "Residential Property Transfer Registration in Fengtai District - Handled in 2019" File 5: "Commercial Real Estate Lease Registration in Dongcheng District - Handled in 2021" File 6: "Industrial Land Transfer Registration in Tongzhou District - Handled in 2017" File 7: "Residential Property Cancellation Registration in Daxing District - Handled in 2014" File 8: "Commercial Real Estate Mortgage Registration in Shijingshan District - Handled in 2022" Perform lexical analysis on the query information of all efficient users and inefficient users, extract keywords, and obtain the set A = {Haidian District, residential, property rights registration, latest, commercial real estate, mortgage registration, Chaoyang District, after 2020, industrial land, alteration registration, house, real estate registration, mortgage, before, registration}; Perform lexical analysis on the classification tags of all files, extract keywords, and obtain the set B = {Haidian District, residential, property rights registration, 2018, Chaoyang District, commercial real estate, mortgage registration, 2016, Xicheng District, industrial land, alteration registration, 2015, Fengtai District, residential property transfer registration, 2019, Dongcheng District, commercial real estate lease registration, 2021, Tongzhou District, industrial land transfer registration, 2017, Daxing District, residential property cancellation registration, 2014, Shijingshan District, commercial real estate mortgage registration, 2022, registration}; Calculate the intersection of set A and set B , and obtain = {Haidian District, residential, property rights registration, commercial real estate, mortgage registration, Chaoyang District, industrial land, alteration registration, registration}, count the number of elements in the intersection = 9, calculate the union of set A and set B , count the number of elements in the union = 32; According to the Jaccard similarity coefficient formula , calculate and obtain ; Assume that the threshold is set to 0.5, compare the calculated matching degree score of 0.28 with the threshold of 0.5. Since 0.28 ≤ 0.5, it indicates that there is a large difference between the query information of efficient users and inefficient users and the classification tags of the files; In step four, the process of determining whether to change the classification tags includes: Calculate the ratio of the actual number of efficient users to the ideal number of efficient users to obtain the ratio of actual efficient users, and compare the ratio of actual efficient users with the threshold value; If the ratio of actual efficient users is greater than or equal to the threshold value, the classification label of the file needs to be changed; If the ratio of actual efficient users is less than the threshold value, the classification label of the file does not need to be changed; Among them, the ideal number of efficient users can be obtained by analyzing the behavior data of users querying files in the past period of time, and statistically calculating the average number of efficient users under the existing classification label as the ideal number of efficient users; Step Five: If the classification label needs to be changed, establish a query information database based on efficient users and inefficient users, extract keywords, and identify the keywords that meet both efficient users and inefficient users as the new classification label; In Step Five, the data included in the query information database includes: Record the user's identity identifier (such as user number, name, etc.) and user type (efficient user or inefficient user); Completely record the original query information such as keywords, phrases or sentences input by the user when querying the real estate registration file; Record the identification information such as the number and name of the real estate registration file involved in the user's query operation; In Step Five, the method of extracting keywords is: Use the lexical analysis tool in natural language processing technology to extract representative keywords after processing the user's query information; Among them, the keyword refers to the words or phrases that are representative and can reflect the main content and core meaning of the query information or classification label after processing the query information of efficient users and inefficient users and the classification label of the file; In Step Five, the process of identifying the keywords that meet both efficient users and inefficient users includes: Label the extracted keywords to clarify whether each keyword belongs to an efficient user or an inefficient user, and organize the extracted keywords and the corresponding user types into the form of a data set; Use the Naive Bayes model to identify the keywords that meet both efficient users and inefficient users. The process includes: Calculate the prior probability. The prior probability refers to the probability of a certain user type appearing without any query information. The calculation method is to count the number of efficient users and inefficient users in the data set, and then divide them by the total number of users respectively. The calculation formula is , ; Among them, P(high) and P(low) are the prior probabilities of efficient users and inefficient users respectively, N 高, N 低 are the numbers of high - efficiency users and low - efficiency users respectively, and N is the total number of high - efficiency users and low - efficiency users; Calculate the conditional probability. The conditional probability refers to the probability that a certain keyword appears given a certain user type. The calculation formula is , ; where, , are the conditional probabilities of high - efficiency users and low - efficiency users respectively, is the keyword, N 高 , N 低 are the numbers of high - efficiency users and low - efficiency users respectively, is the number of records containing the keyword in the query records of high - efficiency users, is the number of records containing the keyword in the query records of low - efficiency users; According to Bayes' theorem, for a keyword , compare ×P(high) and ×P(low); If ×P(high) > ×P(low), the keyword is more likely to be related to high - efficiency users; If ×P(high) < ×P(low), the keyword is more likely to be related to low - efficiency users; If ×P(high) ≈ ×P(low), the keyword fits both high - efficiency users and low - efficiency users; Repeat the prediction step for all keywords in the dataset, and find the keywords that are judged by the model to fit both high - efficiency users and low - efficiency users as new classification labels.

[0021] The technical solution and beneficial points of the embodiments of this application are as follows: Count the number of inquiries for each real estate registration file within the historical period and the total number of real estate files, calculate the relative inquiry ratio of the real estate registration files, divide the real estate registration files into high-frequency inquiry files and low-frequency inquiry files. For high-frequency inquiry files, screen out the users with high-frequency and stable inquiries, establish a corresponding user database, calculate the inquiry efficiency of the high-frequency inquiry files, compare it with the threshold, mark the files with an inquiry efficiency greater than or equal to the threshold as efficient files. For efficient files, divide the corresponding users in the user database into efficient users and inefficient users, and determine whether the difference between the inquiry information of the efficient users and inefficient users and the classification labels of the files is large. If it is large, judge whether it is necessary to change the classification label according to the ratio of the actual number of efficient users to the ideal number of efficient users. If the classification label needs to be changed, establish an inquiry information database based on the efficient users and inefficient users, extract keywords, and identify the keywords that conform to both efficient users and inefficient users as the new classification labels. Through dividing the real estate registration files into high-frequency inquiry files and low-frequency inquiry files, screening out the users with high-frequency and stable inquiries corresponding to the high-frequency inquiry files, calculating the inquiry efficiency of the high-frequency inquiry files, dividing the high-frequency inquiry files into efficient files and inefficient files, dividing the users corresponding to the efficient files into efficient users and inefficient users, analyzing the difference between the inquiry information of the efficient users and inefficient users and the files, and judging whether to change the classification label of the files, and identifying the keywords that conform to both efficient users and inefficient users as the new classification labels, the present application solves the problem of low efficiency in user inquiries due to unreasonable label classification of files during the sorting process of real estate registration files, and improves the efficiency of user inquiries.

[0022] Embodiment 2 Please refer to Figure 2 As shown in the figure, a real estate registration file sorting service system according to an embodiment of the present invention includes: File inquiry frequency classification module: Count the number of inquiries for each real estate registration file within the historical period and the total number of real estate files, calculate the relative inquiry ratio of the real estate registration files, and divide the real estate registration files into high-frequency inquiry files and low-frequency inquiry files; The relative inquiry ratio refers to the ratio of the number of inquiries of real estate registration files to the total number of real estate registration files within the historical period, and the calculation formula is:

[0023] The process of dividing the real estate registration files into high-frequency inquiry files and low-frequency inquiry files includes: Compare the relative inquiry ratio of the real estate registration files with the threshold, divide the real estate registration files with a relative inquiry ratio greater than or equal to the threshold into high-frequency inquiry files, and divide the real estate registration files with a relative inquiry ratio less than the threshold into low-frequency inquiry files; Among them, the setting of the threshold for the relative query ratio includes: calculating the relative query ratio of all real estate registration files within the historical period and calculating the average value, and taking the average value of the relative query ratio as the threshold; High-frequency stable user screening module: For high-frequency query files, screen out users with high-frequency stable queries and establish a corresponding user library; The process of screening out users with high-frequency stable queries is as follows: Divide the historical period into n small time periods. For each user, count the number of queries in each small time period to generate a query count sequence; Compare the number of queries of each user in the small time period with the set threshold. If the number of queries is greater than the threshold in a small time period, mark this small time period as a high-frequency time period; Among them, the setting of the query count threshold includes: calculating the average value of the query count sequence and taking the average value as the query count threshold; Count the number of high-frequency time periods of each user within the historical period, calculate the ratio with the total number of small time periods to obtain the high-frequency time period ratio; Calculate the standard deviation of the query count sequence of each user. The formula is: , where n is the number of small time periods, is the number of queries in each small time period, is the average value of the query counts; For each user, compare the high-frequency time period ratio and the standard deviation of the query count sequence with the threshold respectively, and screen out users with a high-frequency time period ratio greater than or equal to the threshold and a query time standard deviation greater than or equal to the threshold as high-frequency stable query users; Efficient file and user division module: Calculate the query efficiency of high-frequency query files, compare it with the threshold, mark files with a query efficiency greater than or equal to the threshold as efficient files, and divide the corresponding users in the user library into efficient users and inefficient users; The calculation process of the query efficiency of the high-frequency query files includes: Record the start time and end time of each query of the real estate registration file, calculate the time consumption of each query, then add up the time consumptions of all query operations within a certain time period (such as one month, one quarter), and divide by the number of queries to obtain the average query time; Statistically calculate the ratio of the number of times of successfully querying the required file to the total number of queries within a certain time to obtain the query success rate. Among them, a successful query means that the user can accurately obtain the required file information within the specified time; When the user inputs keywords or other retrieval conditions for query, calculate the ratio of the number of files related to the user's actual needs in the retrieval results to the total number of retrieval results to obtain the retrieval result accuracy rate; The query efficiency is comprehensively calculated based on the average query time, query success rate, and retrieval result accuracy. The process includes: Perform data standardization on the average query time, query success rate, and retrieval result accuracy. For the average query time, the smaller the value, the better. Use the formula for standardization to obtain the standardized value of the average query time, where X is the original average query time, Max is the maximum value of the average query time, and Min is the minimum value of the average query time; For the query success rate and retrieval result accuracy, the larger the value, the better. Use the formula for standardization to obtain the standardized value of the query success rate and the standardized value of the retrieval result accuracy, where X is the original query success rate or retrieval result accuracy, Max is the maximum value of the query success rate or retrieval result accuracy, and Min is the minimum value of the query success rate or retrieval result accuracy; Multiply the standardized data by their respective weights and then add them to obtain the comprehensive query efficiency. The formula is: where, 、 、 are the weights of the standardized value of the average query time, the standardized value of the query success rate, and the standardized value of the retrieval result accuracy respectively; Compare the calculated query efficiency of each file with the set threshold. If the query efficiency of the file is higher than the threshold, classify it as a high-efficiency file. If the query efficiency of the file is lower than the threshold, classify it as a low-efficiency file; Among them, the setting of the threshold for file query efficiency includes: adding up the query efficiencies of all files to obtain the total query efficiency, and dividing the total query efficiency by the number of files to obtain the average value of file query efficiency. The calculation formula is: , and use the average value of the query efficiency as the threshold, where is the average value, is the query efficiency of the i-th file, and n is the total number of files; The process of classifying the users corresponding to high-efficiency query files into high-efficiency users and low-efficiency users includes: For each user corresponding to a high-efficiency file, count the number N of high-efficiency files among the files queried by the user in the historical period 高效 and the total number of queries N 总 , and calculate the ratio of high-efficiency files: , where R i is the ratio of high-efficiency files; Compare the ratio of high-efficiency files of each user with the set threshold. If the ratio of high-efficiency files is greater than or equal to the threshold, the user is a high-efficiency user. If the ratio of high-efficiency files is less than the threshold, the user is a low-efficiency user; Among them, the threshold setting of the high-efficiency file occupancy ratio includes: listing the distribution of the high-efficiency file occupancy ratios of all users, and taking the median of the high-efficiency file occupancy ratios as the threshold of the high-efficiency occupancy ratio; Classification label difference evaluation module: judging whether the difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the files is large. If it is large, judging whether the classification labels need to be changed according to the ratio of the actual number of high-efficiency users to the ideal number of high-efficiency users; The process of judging whether the difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the files is large includes: Collect the query information of high-efficiency users and low-efficiency users. Among them, the query information includes keywords, phrases or sentences input by users during queries, and obtain the classification labels of the files; Using the lexical analysis tool in natural language processing technology, process the query information of high-efficiency users and low-efficiency users and the classification labels of the files, extract keywords, and generate the query information keyword set A and the classification label keyword set B respectively; Calculate the intersection of set A and set B, that is, the set composed of elements that belong to both A and B, denoted as , and count the number of elements in the intersection, denoted as ; Calculate the union of set A and set B, that is, the set composed of elements that belong to A or belong to B, denoted as , and count the number of elements in the union, denoted as ; According to the formula of Jaccard similarity coefficient: , calculate the matching degree score between the query information of high-efficiency users and low-efficiency users and the file classification labels; Among them, the Jaccard similarity coefficient is an index used to measure the similarity between two sets, and its value range is [0,1]. The closer the coefficient is to 1, the higher the similarity between the two sets, and the closer it is to 0, the lower the similarity; Compare the matching degree score between the query information of high-efficiency users and low-efficiency users and the file classification labels with the threshold. If the matching degree score is less than or equal to the threshold, it means that the difference between the query information of high-efficiency users and low-efficiency users and the file classification labels is large; The process of judging whether the classification labels need to be changed includes: Calculate the ratio of the actual number of high-efficiency users to the ideal number of high-efficiency users to obtain the actual high-efficiency user occupancy ratio, and compare the actual high-efficiency user occupancy ratio with the threshold; If the actual high-efficiency user occupancy ratio is greater than or equal to the threshold, the classification labels of the files need to be changed; If the actual high-efficiency user occupancy ratio is less than the threshold, the classification labels of the files do not need to be changed; Among them, the number of ideal efficient users can be obtained by analyzing the behavior data of users querying archives in the past period of time, and statistically calculating the average number of efficient users under the existing classification labels as the number of ideal efficient users; New classification label generation module: If the classification label needs to be changed, establish a query information database based on efficient users and inefficient users, extract keywords, and identify the keywords that conform to both efficient users and inefficient users as the new classification label; The data included in the said query information database are: Record the identity identification of users (such as user number, name, etc.), and user type (efficient user or inefficient user); Completely record the original query information such as keywords, phrases or sentences input by users when querying real estate registration archives; Record the identification information such as the number and name of the real estate registration archives involved in the user's query operation; The method for extracting keywords is: Using the lexical analysis tool in natural language processing technology, extract representative keywords after processing the user's query information; Among them, the keyword refers to the words or phrases that are representative and can reflect the main content and core meaning of the query information or classification label after processing the query information of efficient users and inefficient users and the classification labels of the archives; In step five, the process of identifying the keywords that conform to both efficient users and inefficient users includes: Label the extracted keywords to clarify which user type each keyword belongs to, and organize the extracted keywords and the corresponding user types in the form of a data set; Use the Naive Bayes model to identify the keywords that conform to both efficient users and inefficient users. The process includes: Calculate the prior probability. The prior probability refers to the probability of a certain user type appearing without any query information. The calculation method is to count the number of efficient users and inefficient users in the data set, and then divide them by the total number of users respectively. The calculation formula is , ; Among them, P(high) and P(low) are the prior probabilities of efficient users and inefficient users respectively, N 高 , N 低 are the numbers of efficient users and inefficient users respectively, and N is the total number of efficient users and inefficient users; Calculate the conditional probability. The conditional probability refers to the probability of a certain keyword appearing given a certain user type. The calculation formula is , ; Among them, , are the conditional probabilities for highly efficient users and low - efficient users respectively, is the keyword, N 高 , N 低 are the numbers of highly efficient users and low - efficient users respectively, is the number of records containing the keyword in the query records of highly efficient users, is the number of records containing the keyword in the query records of low - efficient users; According to Bayes' theorem, for a keyword , compare ×P(high) and ×P(low); If ×P(high) > ×P(low), the keyword is more likely to be related to highly efficient users; If ×P(high) < ×P(low), the keyword is more likely to be related to low - efficient users; If ×P(high) ≈ ×P(low), the keyword fits both highly efficient users and low - efficient users; Repeat the prediction steps for all keywords in the dataset, and find the keywords that are judged by the model to fit both highly efficient users and low - efficient users as new classification labels.

[0024] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above - mentioned embodiments. What is described in the above - mentioned embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A method for sorting real estate registration files, characterized in that: Including: Step 1: Count the number of queries for each real estate registration file within the historical period and the total number of real estate files, calculate the relative query ratio of the real estate registration files, and divide the real estate registration files into high-frequency query files and low-frequency query files; Step 2: For high-frequency query files, screen out users with high-frequency and stable queries and establish a corresponding user database; Step 3: Calculate the query efficiency of high-frequency query files, compare it with a threshold, mark the files with query efficiency greater than or equal to the threshold as efficient files, and for efficient files, divide the corresponding users in the user database into efficient users and inefficient users; Step 4: Determine whether the difference between the query information of efficient users and inefficient users and the classification labels of the files is large. If it is large, determine whether the classification labels need to be changed according to the ratio of the actual number of efficient users to the ideal number of efficient users; Step 5: If the classification labels need to be changed, establish a query information database based on efficient users and inefficient users, extract keywords, and identify the keywords that meet both efficient users and inefficient users as the new classification labels.

2. A method for sorting real estate registration files according to claim 1, characterized in that: The method of dividing the real estate registration files into high-frequency query files and low-frequency query files is: Compare the relative query ratio of the real estate registration files with a threshold; Those with a relative query ratio of the real estate registration files greater than or equal to the threshold are classified as high-frequency query files; Those with a relative query ratio of the real estate registration files less than the threshold are classified as low-frequency query files.

3. A method for sorting real estate registration files according to claim 2, characterized in that: The calculation method of the relative query ratio of the real estate registration files is: Within the historical period, calculate the ratio of the number of queries of the real estate registration files to the total number of real estate registration files to obtain the relative query ratio of the real estate registration files.

4. A method for sorting real estate registration files according to claim 1, characterized in that: The process of screening out users with high-frequency and stable queries includes: Divide the historical period into n small time periods. For each user, count the number of queries in each small time period to generate a query count sequence; Compare the number of queries of each user in the small time period with a set threshold. If the number of queries is greater than the threshold in a small time period, mark this small time period as a high-frequency time period; Count the number of high-frequency time periods of each user within the historical period, calculate the ratio with the total number of small time periods to obtain the high-frequency time period ratio; For each user, compare the high-frequency time period ratio and the standard deviation of the query count sequence with the threshold respectively, and screen out users with a high-frequency time period ratio greater than or equal to the threshold and a query time standard deviation greater than or equal to the threshold as high-frequency and stable query users.

5. A method for sorting real estate registration files according to claim 1, characterized in that: The calculation process of the query efficiency of the high-frequency query files includes: Record the start time and end time of each query of the real estate registration archives, calculate the time consumption of each query, then add up the time consumptions of all query operations within a certain period of time and divide by the number of queries to obtain the average query time; Statistically calculate the ratio of the number of times the required archives are successfully queried to the total number of queries within a certain period of time to obtain the query success rate; A successful query means that the user can accurately obtain the required archive information within the specified time; When the user enters keywords or other retrieval conditions for query, calculate the ratio of the number of archives related to the user's actual needs in the retrieval results to the total number of retrieval results to obtain the retrieval result accuracy rate; Standardize the average query time, query success rate, and retrieval result accuracy rate. Multiply the standardized data by their respective weights and then add them up to obtain the query efficiency.

6. A method for sorting real estate registration archives according to claim 1, characterized in that: The process of dividing the corresponding users in the user library into high-efficiency users and low-efficiency users includes: For each user corresponding to high-efficiency archives, count the number of times the archives belonging to high-efficiency archives and the total number of queries in the archives queried by the user during the historical period, and calculate the ratio of high-efficiency archives; Compare the ratio of high-efficiency archives of each user with a set threshold; If the ratio of high-efficiency archives is greater than or equal to the threshold, the user is a high-efficiency user; If the ratio of high-efficiency archives is less than the threshold, the user is a low-efficiency user.

7. A method for sorting real estate registration archives according to claim 1, characterized in that: The method for judging whether the difference between the query information of high-efficiency users and low-efficiency users and the classification labels of the archives is large is as follows: Collect the query information of high-efficiency users and low-efficiency users, extract keywords, and generate a set of query information keywords A and a set of classification label keywords B respectively; Calculate the intersection and union of set A and set B respectively, and count the number of elements in the intersection and union; Calculate the matching degree score between the query information of high-efficiency users and low-efficiency users and the archive classification labels according to the formula of Jaccard coefficient; Compare the matching degree score between the query information of high-efficiency users and low-efficiency users and the archive classification labels with a threshold; If the matching degree score is less than or equal to the threshold, it means that the difference between the query information of high-efficiency users and low-efficiency users and the archive classification labels is large.

8. A method for sorting real estate registration archives according to claim 1, characterized in that: The method for identifying keywords that meet both high-efficiency users and low-efficiency users is as follows: Label the extracted keywords to clarify whether each keyword belongs to a high-efficiency user or a low-efficiency user, and organize the extracted keywords and the corresponding user types into the form of a data set; Use the Naive Bayes model to identify keywords that meet both high-efficiency users and low-efficiency users.

9. A method for sorting real estate registration archives according to claim 8, characterized in that: The process of using the Naive Bayes model to identify keywords that meet both high-efficiency users and low-efficiency users includes: Calculate the prior probability and conditional probability of high-efficiency users and low-efficiency users respectively; Compare the product of the prior probability and the conditional probability of high-efficiency users with the product of the prior probability and the conditional probability of low-efficiency users; If the product of the prior probability and the conditional probability of high-efficiency users is approximately equal to the product of the prior probability and the conditional probability of low-efficiency users, it indicates that the keyword conforms to both high-efficiency users and low-efficiency users.

10. A real estate registration file sorting service system, characterized in that: File query frequency classification module: Count the number of queries of each real estate registration file in the historical period and the total number of real estate files, calculate the relative query ratio of the real estate registration file, and divide the real estate registration file into high-frequency query files and low-frequency query files; High-frequency stable user screening module: For high-frequency query files, screen out users with high-frequency stable queries and establish a corresponding user library; Efficient file and user classification module: Calculate the query efficiency of high-frequency query files, compare it with the threshold, mark the files with query efficiency greater than or equal to the threshold as efficient files, and for efficient files, divide the corresponding users in the user library into high-efficiency users and low-efficiency users; Classification label difference evaluation module: Judge whether the difference between the query information of high-efficiency users and low-efficiency users and the classification label of the file is large. If it is large, judge whether it is necessary to change the classification label according to the ratio of the actual number of high-efficiency users to the ideal number of high-efficiency users; New classification label generation module: If the classification label needs to be changed, establish a query information library based on high-efficiency users and low-efficiency users, extract keywords, and identify the keywords that conform to both high-efficiency users and low-efficiency users as the new classification label.

Citation Information

Patent Citations

  • Search optimization method and device, electronic equipment, storage medium and program product

    CN119025619A

  • Archive information storage method and system

    CN120045516A

  • Data indexing system using dynamic tags

    US20200342014A1