Technical achievement classification and retrieval method and system
By performing word frequency analysis and keyword domain mapping of the metadata of the scientific and technological literature database, optimizing the literature classification and indexing structure, and combining user feedback to revise the parameters, the problems of inefficient search efficiency and inaccurate results in the face of data volume growth are solved, and efficient and accurate literature search and resource management are achieved.
Patent Information
- Application Number
- CN202510225223.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When the existing technology faces the rapid growth of data volume, it is difficult to respond quickly and accurately to users' complex and changing query needs, and lacks effective user feedback mechanisms and adaptive optimization measures, resulting in inefficient retrieval efficiency and inaccurate results, which affects scientific research efficiency and rational utilization of resources.
Through word frequency analysis based on the metadata of the scientific and technological literature database, keyword field mapping tables are generated, literature classification process is optimized, literature index structure is updated, classification and index parameters are revised through user feedback, and search interfaces that integrate search engine functions are developed to realize search logic programming for keywords and fields combinations.
It significantly improves the classification accuracy and search efficiency of the literature, improves user satisfaction and system flexibility, saves scientific researchers' time in finding and analyzing scientific data, and optimizes the allocation of R&D resources.
Smart Images

Figure CN120162353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and particularly to a method and system for classifying and retrieving technical achievements. Background Art
[0002] The technical field of information retrieval involves the development and application of algorithms, systems, and software tools for quickly and effectively retrieving information from large datasets. The core goal of the field is to optimize the retrieval process to ensure highly relevant, efficient, and accurate search results. Information retrieval is widely used in web search engines, online databases, digital libraries, and fields that need to process and access large amounts of documents and data. With the continuous increase in data volume, information retrieval technology is gradually integrating more artificial intelligence and machine learning methods to improve search algorithms and enhance the user experience.
[0003] Among them, the method for classifying and retrieving technical achievements refers to a technical means for classifying and quickly retrieving scientific and technological achievements. This method enables users to effectively manage and access scientific and technological achievements, and is particularly suitable for the technical management departments of scientific research institutions, universities, and enterprises. Using this method, the filing, retrieval, and reuse of scientific and technological achievements can be accelerated, significantly improving the efficiency of current scientific and technological innovation. Through precise classification and efficient retrieval, it helps users save time in searching for and analyzing scientific and technological materials, and optimizes the allocation of R & D resources.
[0004] Although the prior art can handle large datasets and provide relatively effective retrieval results, it still shows certain limitations in document management and quick retrieval. Traditional information retrieval relies on relatively fixed index structures and retrieval algorithms, which limits its adaptability and flexibility in the face of rapidly growing data volumes. With the development of technology, the types and structures of data are constantly evolving, and the old indexes are difficult to quickly and accurately respond to the complex and changing query requirements of users. The lack of an effective user feedback mechanism and adaptive optimization measures makes it difficult for the prior art to quickly respond to the specific needs of users in actual operation. Users will encounter problems such as low retrieval efficiency and inaccurate results when conducting scientific and technological literature retrieval, which affects scientific research efficiency and the rational utilization of resources. Summary of the Invention
[0005] The object of the present invention is to solve the drawbacks existing in the prior art, and to propose a method and system for classifying and retrieving technical achievements.
[0006] To achieve the above object, the present invention adopts the following technical solution: A method for classifying and retrieving technical achievements, comprising the following steps:
[0007] S1: Based on the metadata of the documents in the scientific and technological literature database, perform word frequency analysis, record the frequency distribution of keywords, associate them with the document metadata, generate the metadata association analysis results, and map the keywords to classification labels according to the metadata association analysis results to obtain the keyword domain mapping table;
[0008] S2: Based on the keyword domain mapping table, scan the keywords of the documents, assign the corresponding domain labels, obtain the preliminary classification results, evaluate the accuracy of the classification according to the preliminary classification results, optimize the classification parameters, and obtain the optimized classification results;
[0009] S3: Based on the optimized classification results, update the document index information in the database to obtain the new document index architecture, implement the index update plan according to the new document index architecture, analyze the user feedback, revise the classification and index parameters, and obtain the maintenance and update strategy;
[0010] S4: Based on the maintenance and update strategy, develop a document retrieval interface, integrate the search engine function, implement the search logic programming of the keyword and field combination, obtain the prototype of the retrieval interface, perform query optimization according to the prototype of the retrieval interface, optimize the search engine configuration by analyzing the query log, and generate the optimized classification retrieval scheme.
[0011] As a further solution of the present invention, the specific steps for obtaining the keyword domain mapping table are as follows:
[0012] S111: Based on the metadata of the documents in the scientific and technological literature database, including keywords and document classification labels, capture the document metadata from different sources, screen and sort out the duplicate keywords and classification labels, and establish the basis for document metadata analysis;
[0013] S112: Based on the basis for document metadata analysis, perform word frequency statistics operations, analyze and record the number of occurrences of each keyword to obtain the keyword frequency distribution data;
[0014] S113: According to the keyword frequency distribution data, conduct the correlation analysis between keywords and document classification labels, calculate the correlation weight between keywords and classification labels, through the formula:
[0015]
[0016] Generate the keyword domain mapping table, where W a represents the correlation weight between the target keyword and the document classification label, f a represents the frequency of keyword a appearing in the document, C a represents the classification label frequency corresponding to keyword a, and ∑f a is the total sum of the keyword occurrence frequencies.
[0017] As a further solution of the present invention, the steps for obtaining the preliminary classification result are specifically as follows:
[0018] S211: Based on the keyword-domain mapping table, access the literature database, extract the keyword list of each literature, including database query, data extraction, and preliminary data processing, and generate a literature keyword list;
[0019] S212: Based on the literature keyword list, determine the corresponding domain label for each keyword, capture the corresponding domain label, mark the unclassified keywords as unclassified, and generate the mapping result from keywords to domains;
[0020] S213: Summarize the mapping results from keywords to domains of each literature, adopt a weighted statistical method to optimize the classification accuracy, through the formula:
[0021]
[0022] By calculating the weighted scores of each domain, obtain the preliminary classification result, where δ b,k is the indicator function of keyword k, w k is the weight of the keyword, N is the total number of keywords in the literature, and L b represents the weighted score of the target domain.
[0023] As a further solution of the present invention, the steps for obtaining the optimized classification result are specifically as follows:
[0024] S221: Based on the preliminary classification result, extract literature data from the database, including literature identifiers, keyword lists, and preliminary domain labels, and evaluate the classification of the literature data to obtain the classification accuracy evaluation result;
[0025] S222: According to the classification accuracy evaluation result, by comparing the preliminary domain label of each literature with the truly labeled domain, calculate the number of successful and failed classification instances, and obtain the accuracy of the domain label;
[0026] S223: According to the accuracy of the domain label, adjust the parameters in the classification algorithm, including weights and thresholds, optimize the overall performance of the classification, and adopt the formula:
[0027]
[0028] Generate the optimized classification result, where P c represents the prediction accuracy of the optimized domain, f c is the preliminary classification accuracy of the domain, s c is an additional parameter adjusted based on the domain characteristics, and λ is used to adjust f c and sc The regularization coefficient of the contribution.
[0029] As a further solution of the present invention, the acquisition steps of the new literature index architecture are specifically as follows:
[0030] S311: Based on the optimized classification results, call the associated literature records from the database, including the unique identifier, keywords, original classification, and new classification of each literature, and use the database query command to determine the integrity and accuracy of the data, and generate a literature record list;
[0031] S312: Apply the current index structure update algorithm to the literature record list, through the formula:
[0032] I d = hash(K d + C old + C new )
[0033] Calculate the new index key value of each literature, and generate the current index mapping. Among them, I d is the new index, K d is the literature keyword, C old and C new are the old classification and the new classification respectively;
[0034] S313: Update the current index mapping to the database, including replacing the current index in the database and adding the new index, and generate a new literature index architecture.
[0035] As a further solution of the present invention, the acquisition steps of the maintenance update strategy are specifically as follows:
[0036] S321: Based on the new literature index architecture, design and implement an index update plan according to the current literature review and index structure, and generate updated index parameters by comparing and analyzing the differences between the current index and the new requirements;
[0037] S322: Use the updated index parameters, combined with the real-time user feedback data set, to check the data integrity and accuracy of each index item, and generate a difference analysis result;
[0038] S323: According to the difference analysis result, revise the index items with insufficient classification effect and user satisfaction, and apply the formula:
[0039]
[0040] Optimize the overall deviation of the revised index item to generate a maintenance update strategy, where p h represents the original index parameter, q h represents the revised index parameter, wh r represents the weight coefficient, M represents the total number of parameters, and E represents the weighted mean square error evaluation value of the revision effect.
[0041] As a further solution of the present invention, the steps for obtaining the optimized classification retrieval solution are specifically as follows:
[0042] S411: Based on the maintenance and update strategy, develop the current literature retrieval interface, integrate search engine technology, perform combined search of keywords and multiple fields, and generate a fully functional retrieval interface prototype;
[0043] S412: Conduct real-time operation simulation on the fully functional retrieval interface prototype, capture preliminary user feedback and query logs, analyze the data and identify common query problems and user preferences, adjust the configuration parameters of the search engine, optimize the user experience, and generate a preliminarily optimized search engine configuration;
[0044] S413: Based on the preliminarily optimized search engine configuration, apply the formula:
[0045]
[0046] Adjust multiple parameters, optimize the query efficiency and accuracy by controlling the weight distribution, and generate an optimized classification retrieval solution, where r g represents the real-time response time, t g represents the target response time, w g is the weight set for different query types, Q represents the optimized total error value, and G represents the total number of query types.
[0047] A technical achievement classification and retrieval system, which is used to execute the above technical achievement classification and retrieval method, and the system includes:
[0048] The keyword analysis module, based on the metadata of the documents in the scientific and technological literature database, identifies and records the keywords of each document, performs frequency statistics on the keywords, sorts the keywords according to the statistical data, calculates the frequency distribution of the keyword occurrences, and generates a keyword frequency distribution table;
[0049] The keyword mapping module, based on the keyword frequency distribution table, analyzes the relationship between the keywords and the domain labels, performs domain label mapping on the keywords, associates the keywords with the target domain labels, and constructs a keyword-domain mapping table;
[0050] The classification annotation module scans the keywords in the documents through the keyword-domain mapping table, annotates the domain labels of each document, uses the labels to evaluate the preliminary accuracy of the classification, adjusts and optimizes the classification annotation parameters, and obtains an optimized classification result;
[0051] The index update module reconstructs the literature index structure in the database according to the optimized classification result, performs the update operation of the index, analyzes the user feedback information, adjusts the index and classification parameters according to the feedback, and establishes a maintenance and update strategy.
[0052] The retrieval interface development module uses the maintenance and update strategy to develop the current literature retrieval interface, integrates the search function of the combination of keywords and literature attribute fields, tests and optimizes the search logic, analyzes through the query log, adjusts the configuration parameters of the search engine, and generates an optimized classification retrieval scheme.
[0053] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0054] In the present invention, through the word frequency analysis of the literature metadata, the keyword frequency record and the correlation analysis of the metadata, the accurate mapping from keywords to domain labels is realized. This mapping method optimizes the literature classification process, significantly improves the accuracy of the preliminary classification. Further, through the evaluation of the classification accuracy rate and parameter optimization, the efficiency and adaptability of the classification result are ensured, so as to realize more effective management and access to scientific and technological achievements. By updating the literature index architecture and implementing the index update plan, not only the filing and retrieval speed of scientific and technological achievements are accelerated, but also the classification and index parameters are revised through user feedback, further improving the flexibility of the retrieval system and user satisfaction. This comprehensive optimization strategy provides strong support for the literature retrieval interface. Through the logical programming of the combination of keywords and fields, the retrieval is more accurate and fast, thus greatly saving the time of scientific research personnel in searching and analyzing scientific and technological materials and optimizing the allocation of R & D resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic diagram of the working process of the present invention;
[0056] Figure 2 It is a flowchart of the keyword domain mapping table in the present invention;
[0057] Figure 3 It is a flowchart of the preliminary classification result in the present invention;
[0058] Figure 4 It is a flowchart of the optimized classification result in the present invention;
[0059] Figure 5 It is a flowchart of the new literature index architecture in the present invention;
[0060] Figure 6 It is a flowchart of the maintenance and update strategy in the present invention;
[0061] Figure 7 It is a flowchart of the optimized classification retrieval scheme in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0062] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0064] Embodiment:
[0065] Please refer to Figure 1 , the present invention provides a technical solution: a method for classifying and retrieving technical achievements, including the following steps:
[0066] S1: Based on the metadata of the documents in the scientific and technological literature database, perform word frequency analysis, record the keyword frequency distribution, and associate it with the document metadata to generate a metadata association analysis result. According to the metadata association analysis result, map the keywords to classification labels to obtain a keyword domain mapping table;
[0067] S2: Based on the keyword domain mapping table, scan the keywords of the document, assign the corresponding domain labels, obtain a preliminary classification result, evaluate the accuracy rate of the classification according to the preliminary classification result, optimize the classification parameters, and obtain an optimized classification result;
[0068] S3: Based on the optimized classification result, update the document index information in the database to obtain a new document index architecture. According to the new document index architecture, implement an index update plan, analyze user feedback, revise the classification and index parameters, and obtain a maintenance and update strategy;
[0069] S4: Based on the maintenance and update strategy, develop a document retrieval interface, integrate the search engine function, implement the search logic programming of the combination of keywords and fields, obtain a prototype of the retrieval interface, perform query optimization according to the prototype of the retrieval interface, analyze the query log, optimize the search engine configuration, and generate an optimized classification retrieval scheme.
[0070] The results of metadata correlation analysis include statistical charts, trend analysis, and correlation degree scoring. The keyword-domain mapping table includes a domain thesaurus, mapping rules, and application frequencies. The preliminary classification results include label assignment, domain determination, and preliminary error assessment. The optimized classification results include accuracy improvement, error reduction, and parameter optimization records. The new architecture of the literature index includes index levels, index algorithms, and index library structures. The maintenance and update strategy includes the strategy revision cycle, monitoring metrics, and feedback adjustment processes. The prototype of the retrieval interface includes a design blueprint, interaction logic, and functional interfaces. The optimized classification and retrieval scheme includes algorithm efficiency, retrieval accuracy, and system stability.
[0071] Please refer to Figure 2 , and the steps for obtaining the keyword-domain mapping table are specifically as follows:
[0072] S111: Based on the metadata of the literature in the scientific and technological literature database, including keywords and literature classification labels, capture the literature metadata from different sources, screen and sort out duplicate keywords and classification labels, and establish the basis for literature metadata analysis;
[0073] Extract the metadata of the literature from the scientific and technological literature database, which involves accessing the database, querying the required fields, and extracting data, including keywords and literature classification labels, to form an initial keyword-classification label comparison set, which establishes a basic data set for subsequent analysis. In the subsequent data processing, it is necessary to perform quality control on the extracted data, exclude errors or missing values in the data, and ensure the accuracy and integrity of the keywords and classification labels, thereby ensuring the reliability of the data set. The data cleaning process includes identifying abnormal data, handling missing values, and format standardization, etc. The initial set will be used to analyze the classification trends of the literature and the distribution characteristics of the keywords in the future, so as to generate the basis for literature metadata analysis.
[0074] S112: Based on the basis of literature metadata analysis, perform a word frequency statistics operation, analyze and record the number of occurrences of each keyword, and obtain keyword frequency distribution data;
[0075] Perform a word frequency statistics operation on the initial keyword-classification label comparison set. This operation involves the use of statistical analysis software such as Python or R. By programming, calculate the number of occurrences of each keyword. The statistical process requires writing scripts for automated processing, which involves data reading, loop traversal, and frequency calculation, etc. Among them, the word frequency statistics is not only data, but also needs to optimize the performance through appropriate data structures such as dictionaries or hash tables to speed up the processing speed and ensure the efficiency of the statistics. Further obtain keyword frequency distribution data, which provides key raw inputs for subsequent analysis and is a prerequisite for establishing an efficient and accurate literature keyword analysis model. This step provides the necessary data basis for subsequent analysis.
[0076] S113: Based on the keyword frequency distribution data, perform the correlation analysis between keywords and literature classification tags, calculate the correlation weight between keywords and classification tags, through the formula:
[0077]
[0078] Generate a keyword domain mapping table, where W a represents the correlation weight between the target keyword and the literature classification tag, and f a represents the frequency of keyword a appearing in the literature, and C a represents the classification tag frequency corresponding to keyword a, and ∑f a is the total sum of the keyword appearance frequencies;
[0079] The benefit of the formula is to calculate the weight corresponding to each keyword by correlating the keyword frequency and the classification tag frequency, providing a method to quantify the importance of keywords in different classifications, thereby assisting in analyzing the distribution and influence of keywords in various fields; formula detailed explanation and formula calculation derivation process: Assume that the keyword "data analysis" appears in 5 articles, among which 3 articles belong to the "machine learning" classification and 2 articles belong to the "big data" classification. The classification frequencies of "machine learning" and "big data" are 30 times and 20 times respectively. Calculate the weight of "data analysis";
[0080] Data analysis f a = 5, machine learning C1 = 30, big data C2 = 20, and the sum of all keywords ∑f a = 100. Substitute the specific values into the formula:
[0081]
[0082] The result shows that the weight of "data analysis" is 2.5, indicating that among all keywords, "data analysis" has a relatively high distribution weight in the two classifications of "machine learning" and "big data", indicating its relative importance in these two fields.
[0083] Please refer to Figure 3 for the specific steps to obtain the preliminary classification result:
[0084] S211: Based on the keyword domain mapping table, access the literature database, extract the keyword list of each literature, including database query, data extraction, and preliminary data processing, and generate a literature keyword list;
[0085] Access the literature database through an automated script. First, execute the database query statement to extract the keyword list for each piece of literature, including SQL queries and data extraction operations, ensuring the accuracy and integrity of the extracted keyword data. Through a preliminary data cleaning process, exclude any errors or outliers, such as removing illegal characters and correcting formatting errors, to ensure the accuracy of the keyword data list, thereby providing a reliable basis for subsequent analysis and mapping, generating a verified and cleaned literature keyword list, and preparing for the next step of domain label matching.
[0086] S212: Based on the literature keyword list, determine the corresponding domain label for each keyword, capture the corresponding domain label, mark the unclassified keywords as unclassified, and generate the mapping result from keywords to domains;
[0087] Traverse the keyword set of each piece of literature, and find the corresponding domain label for each keyword through the established keyword-domain mapping table, including querying from keywords to domain labels in the mapping table. Each query operation involves the use of a hash table to speed up the keyword search. If a keyword does not have a corresponding domain label in the mapping table, mark it as unclassified, generate the mapping result list from keywords to domains, ensure that each keyword is properly classified or marked, and provide the necessary input data for the comprehensive domain label determination of each piece of literature.
[0088] S213: Aggregate the mapping results from keywords to domains for each piece of literature, adopt a weighted statistical method to optimize the classification accuracy, through the formula:
[0089]
[0090] By calculating the weighted score of each domain, obtain the preliminary classification result, where δ b,k is the indicator function of keyword k, w k is the weight of the keyword, N is the total number of keywords in the literature, and L b represents the weighted score of the target domain;
[0091] The benefit of the formula is that it allows assigning a weight to each keyword in the literature, and this weight reflects the importance of the keyword in the literature, making the classification result not only based on the frequency of keyword occurrence but also its relative importance in the literature, thereby improving the accuracy and relevance of the classification;
[0092] Suppose there are keywords "data analysis" and "machine learning" in a piece of literature. The frequency of the keyword "data analysis" in the "technology" domain is 3 times, and the weight is 0.5. The frequency of "machine learning" in the same domain is 2 times, and the weight is 0.7. The total number of keywords N in the literature = 2:
[0093]
[0094] The result shows that the comprehensive weight in the "technology" field is 1.45, indicating a relatively high degree of association between this document and the "technology" field. A score can be generated for each field based on such calculations, and finally, the field with the highest score is labeled for the document.
[0095] Please refer to Figure 4 , and the specific steps for obtaining the optimized classification result are as follows:
[0096] S221: Based on the preliminary classification result, extract literature data from the database, including literature identifiers, keyword lists, and preliminary field labels, and evaluate the classification of the literature data to obtain the classification accuracy evaluation result;
[0097] Extract literature information from the database to obtain all data related to the preliminary classification, including literature identifiers and keyword lists. The data is the basis for accuracy evaluation. First, perform data cleaning and format unification on the preliminary field labels of each literature, and then create an index table to correspond to the relationship between each label and the literature. The index table will be used for subsequent classification accuracy calculation. Then, calculate the accuracy of each field label based on the data, including calculating the number of successful and failed classification instances. For each field, the success rate and failure rate of its classification should be recorded in detail for the next step of parameter optimization, which is a key link to ensure sufficient data support for subsequent operations.
[0098] S222: According to the classification accuracy evaluation result, calculate the number of successful and failed classification instances by comparing the preliminary field labels of each literature with the true labeled fields to obtain the accuracy of the field labels;
[0099] According to the classification accuracy evaluation results, by comparing the preliminary domain labels of each document with the true annotations, the number of successful and failed classification instances is calculated. First, a dataset is established, which contains the records of thousands of documents. Each record includes the title, abstract, author information, publication date, and preliminary domain labels of the document. The preliminary domain labels are automatically assigned by a preliminary classification system based on the content of the document and include fields such as technology, economy, and medicine. The expert group will review the true domains of these documents and evaluate them against the preliminary domain labels, marking the true domain labels of each document, which is completed by manually reading and analyzing the content of the document to ensure the accuracy and reliability of the annotation. By comparing the preliminary domain labels assigned by the system with the true domain labels of the expert group, the number of successful classification instances (i.e., the cases where the system label is consistent with the true label) and the number of failed classification instances (i.e., the cases where the system label is inconsistent with the true label) are counted. After completing the instance statistics, a simple quantitative comparison method is used to calculate the accuracy of each domain label. For example, the accuracy of the technology field is calculated as the ratio of the number of documents with correct classification of the technology field label to the total number of documents in the technology field. This evaluation method intuitively reflects the annotation efficiency of the preliminary classification system in each field and obtains the accuracy of the domain label.
[0100] S223: According to the accuracy of the domain label, adjust the parameters in the classification algorithm, including weights and thresholds, to optimize the overall performance of the classification. Use the formula:
[0101]
[0102] Generate the optimized classification results, where P c represents the prediction accuracy of the optimized domain, and f c is the preliminary classification accuracy of the domain, s c is an additional parameter adjusted based on the characteristics of the domain, and λ is a regularization coefficient used to adjust the contributions of f c and s c ;
[0103] The benefit of the formula is to optimize the classification results by adjusting the classification parameters, where f c is the preliminary classification accuracy of domain c, directly obtained from data analysis to ensure reflecting the true performance, s c is set based on the experience of domain experts to adjust the strictness of the internal classification criteria in the domain, and λ is determined by cross-validation to make the parameter adjustment more accurate and effective;
[0104] Set f c = 0.85, indicating that the preliminary classification accuracy of a domain is 85%, and s c= 0.15, representing an additional adjustment parameter according to the domain characteristics, λ = 0.5, which is the regularization coefficient used to balance f c and s c 's contributions. The calculation process is as follows:
[0105]
[0106] The result shows that by adjusting the weights and thresholds, the classification accuracy can be improved to 61.67%, which significantly optimizes the performance of the retrieval method.
[0107] Please refer to Figure 5 , and the specific steps for obtaining the new architecture of the literature index are as follows:
[0108] S311: Based on the optimized classification results, call the associated literature records from the database, including the unique identifier, keywords, original classification, and new classification of each literature. Use the database query command to determine the integrity and accuracy of the data, and generate a list of literature records;
[0109] Call all relevant literature records in the database, which include the unique identifier, keywords, original classification, and current classification information of each literature. Ensure that the data extracted from the database is complete and accurate through precise database query statements. This process not only involves data reading but also includes data format verification and preliminary data cleaning. For example, remove those data entries marked incorrectly due to classification errors and merge those records that have been repeatedly entered due to historical reasons. The result of this step is a comprehensive list of current literature records to be updated, which is the basis for the next data processing.
[0110] S312: Apply the current index structure update algorithm to the list of literature records. Through the formula:
[0111] I d = hash(K d + C old + C new )
[0112] Calculate the new index key value for each literature, and generate the current index mapping. Among them, I d is the new index, K d is the literature keyword, C old and C new are the old classification and new classification respectively;
[0113] The advantage of the formula is that by combining keyword and classification information to generate a unique index, it enhances the speed and accuracy of database queries, thus optimizing the information retrieval process;
[0114] For each literature record in the data list obtained in the previous step, the keyword Kd and the original classification C old and the new classification C new are input into a hash function, and a new index I for each document is generated by calculating the hash value d ;
[0115] The specific calculation process is as follows: First, the keywords, original classification, and new classification strings are concatenated, and then the resulting string is input into the hash function to generate an integer hash value as the new index. For example, for the keyword "DataScience", the original classification "CS", and the new classification "Data", the concatenated string is "DataScienceCSData", and its hash value is 23456. This hash value serves as the new index ID for the document. This result indicates that the new index structure can effectively support faster data retrieval and more efficient database management.
[0116] S313: Update the current index mapping in the database, including replacing the current index in the database and adding the new index, to generate a new architecture for the document index;
[0117] Applying the new index mapping to the database involves complex database management operations, such as replacing the index and adding the new index, including write operations to the database, ensuring that all modifications are correctly applied and can be persistently stored, including using transaction processing to ensure the atomicity of operations and avoiding any inconsistencies in the data during the update process. For example, we will set up a transaction log to record each step of the modification so that we can restore to the state before the modification in case of a failure. We also need to update the backup and recovery strategies of the database to adapt to the current index structure. This ensures that even in case of a system failure, the data will not be lost and the normal service state can be quickly restored.
[0118] Please refer to Figure 6 for the specific steps to obtain the maintenance update strategy:
[0119] S321: Based on the new architecture of the document index, design and implement an index update plan according to the current literature review and index structure. By comparing and analyzing the differences between the current index and the new requirements, generate updated index parameters;
[0120] Based on the design of the latest literature review and the existing index structure, by deeply analyzing the index update plan proposed in the literature, evaluating the performance and requirement differences of the existing index system, and identifying the key update points, a systematic update plan is designed. This plan ensures more current effectiveness by carefully comparing and analyzing the differences between the existing index and the new requirements, so that the updated index parameters can better meet the user's needs and improve the response speed and accuracy of the system in practical applications.
[0121] S322: Using the updated index parameters, combined with the real-time user feedback dataset, conduct data integrity and accuracy checks on each index item to generate a differential analysis result;
[0122] Using the updated index parameters, combined with the real-time user feedback dataset, conduct data integrity and accuracy checks on each index item. By specific calculations, obtain the deviation values between each index item and user requirements to ensure that each index can accurately reflect user needs. Through this detailed analysis and calculation process, it is possible to identify which index items most need to be optimized, and then adjust the corresponding parameters to achieve the optimal configuration. This detailed differential analysis helps to precisely locate the problem and propose targeted solution strategies, generating a differential analysis result.
[0123] S323: According to the differential analysis result, revise the index items with insufficient classification effect and user satisfaction, and apply the formula:
[0124]
[0125] Optimize the overall deviation of the revised index items to generate a maintenance and update strategy, where p h represents the original index parameter, q h represents the revised index parameter, w h represents the weight coefficient, M represents the total number of parameters, and E is the weighted mean square error evaluation value of the revision effect;
[0126] The benefit of the formula is that by introducing the weight coefficient w h , it is possible to conduct differential calculations based on the criticality of each parameter, enabling targeted adjustment of the most critical parameters when revising the index parameters, thereby optimizing the performance of the entire index structure;
[0127] Suppose there are three parameters, the original index parameters p1, p2, p3, which are 0.2, 0.5, 0.3 respectively, the revised index parameters q1, q2, q3, which are 0.3, 0.6, 0.4 respectively, and the weight coefficients w1, w2, w3, which are 1, 2, 1.5 respectively. Calculate the square of the difference of each parameter multiplied by the weight to obtain:
[0128] w1·|p1 - q1| 2 = 1·(0.2 - 0.3) 2 = 0.01
[0129] w2·|p2 - q2| 2 = 2·(0.5 - 0.6) 2 = 0.02
[0130] w3·|p3 - q3| 2 = 1.5·(0.3 - 0.4) 2= 0.015
[0131] The sum is: 0.01 + 0.02 + 0.015 = 0.045, and the average is This result indicates that the weighted mean square error evaluation value for the overall parameter revision is 0.015, meaning that the difference between the index structure after parameter adjustment and the original structure is small, and the revision strategy effectively optimizes the index parameters.
[0132] Please refer to Figure 7 , and the specific steps for obtaining the optimized classification retrieval scheme are as follows:
[0133] S411: Based on the maintenance and update strategy, develop the current literature retrieval interface, integrate search engine technology, perform combined search of keywords and multiple fields, and generate a complete functional retrieval interface prototype;
[0134] The development process starts with synthesizing existing maintenance and update strategies, selecting a matching search engine technology, collaborating with the development team to establish the search logic for keyword and field combinations, and through several iterations of development, ensuring the consistency of the basic functions of the retrieval interface and the user interaction design. Integrating elements ensures that the retrieval interface can effectively handle user queries, reflecting the practicality and functionality of the initial prototype of the system design, and generating a retrieval interface prototype that meets expectations.
[0135] S412: Conduct real-time operation simulation on the complete functional retrieval interface prototype, capture preliminary user feedback and query logs, analyze the data and identify common query problems and user preferences, adjust the configuration parameters of the search engine, optimize the user experience, and generate a preliminarily optimized search engine configuration;
[0136] Use the newly developed retrieval interface prototype to organize internal team members and external volunteers for testing, collect feedback and query logs during use, record in detail the system responses and user satisfaction in various query situations, identify key configuration parameters that need to be optimized by analyzing the actual usage data, and make necessary adjustments to the settings of the search engine to ensure that the optimization measures can specifically solve the specific problems encountered by users during the query process, generating a preliminarily optimized search engine configuration.
[0137] S413: Based on the preliminarily optimized search engine configuration, apply the formula:
[0138]
[0139] Adjust multiple parameters, optimize query efficiency and accuracy by controlling weight distribution, and generate an optimized classification retrieval scheme, where r g represents the real-time response time, t g represents the target response time, w gThe weight set for different query types, where Q represents the optimized total error value and G represents the total number of query types;
[0140] The advantage of the formula is that by introducing the query type weight w g , it is possible to customize the response time optimization for different types of queries, which not only improves the efficiency of the search engine but also optimizes the user experience;
[0141] There are three types of queries set, with the actual response times r1, r2, r3 being 0.2 seconds, 0.5 seconds, and 0.3 seconds respectively, and the target response times t1, t2, t3 being 0.1 seconds, 0.4 seconds, and 0.2 seconds respectively. The weights w1, w2, w3 are 1, 2, and 1.5 respectively. Calculate the absolute value of the time difference after weight adjustment to get:
[0142] w1·|r1 - t1| = 1·|0.2 - 0.1| = 0.1
[0143] w2·|r2 - t2| = 2·|0.5 - 0.4| = 0.2
[0144] w3·|r3 - t3| = 1.5·|0.3 - 0.2| = 0.15
[0145] Sum them up to get: Q = 0.1 + 0.2 + 0.15 = 0.45. This result shows that by precisely controlling the weight distribution and response time, the optimized search engine can maintain efficient query processing while significantly improving user satisfaction, and further obtain a finely tuned classification and retrieval scheme.
[0146] A technical achievement classification and retrieval system, which is used to execute the above-mentioned technical achievement classification and retrieval method. The system includes:
[0147] The keyword analysis module, based on the metadata of the documents in the scientific and technological literature database, identifies and records the keywords of each document, performs frequency statistics on the keywords, sorts the keywords according to the statistical data, calculates the frequency distribution of the keywords, and generates a keyword frequency distribution table;
[0148] The keyword mapping module, based on the keyword frequency distribution table, analyzes the relationship between keywords and domain labels, performs domain label mapping on the keywords, associates the keywords with the target domain labels, and constructs a keyword-domain mapping table;
[0149] The classification annotation module scans the keywords in the documents through the keyword-domain mapping table, annotates the domain labels of each document, evaluates the preliminary accuracy of the classification using the labels, adjusts and optimizes the classification annotation parameters, and obtains the optimized classification result;
[0150] The index update module reconstructs the literature index structure in the database according to the optimized classification results, performs the index update operation, analyzes the user feedback information, adjusts the index and classification parameters according to the feedback, and establishes a maintenance and update strategy.
[0151] The retrieval interface development module uses the maintenance and update strategy to develop the current literature retrieval interface, integrates the search function combining keywords and literature attribute fields, tests and optimizes the search logic, analyzes through the query log, adjusts the configuration parameters of the search engine, and generates an optimized classification retrieval scheme.
[0152] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for classifying and retrieving technical achievements, characterized in that: The following steps are involved: Based on the metadata of the documents in the scientific literature database, perform word frequency analysis, record the frequency distribution of keywords, and associate them with the document metadata to generate metadata association analysis results. According to the metadata association analysis results, map the keywords with the classification labels to obtain a keyword field mapping table; Based on the keyword-domain mapping table, the keywords of the document are scanned and labeled with corresponding domain labels to obtain preliminary classification results. According to the preliminary classification results, the accuracy of classification is evaluated, and the classification parameters are optimized to obtain optimized classification results. Based on the optimized classification results, updating the document index information in the database, obtaining a new document index architecture, implementing an index update plan according to the new document index architecture, analyzing user feedback, revising classification and index parameters, and obtaining a maintenance update strategy; Based on the maintenance and update strategy, a document retrieval interface is developed, the search engine function is integrated, and the search logic programming of the keyword and field combination is implemented to obtain a retrieval interface prototype. According to the retrieval interface prototype, query optimization is performed, and by analyzing the query log, the search engine configuration is optimized to generate a tuned classified retrieval plan.
2. The method for classifying and retrieving technical achievements according to claim 1, characterized in that: The steps for obtaining the keyword domain mapping table are specifically as follows: Based on the metadata of documents in the scientific literature database, including keywords and document classification tags, we capture the metadata of documents from different sources, screen and sort out repeated keywords and classification tags, and establish a foundation for document metadata analysis; Based on the document metadata analysis, perform word frequency statistics, analyze and record the number of occurrences of each keyword, and obtain keyword frequency distribution data; According to the keyword frequency distribution data, the correlation analysis between keywords and document classification labels is performed, and the correlation weight between keywords and classification labels is calculated through the formula: Generate a keyword domain mapping table, where W a represents the association weight between the target keyword and the document classification label, f a represents the frequency of keyword a in the literature, C a represents the frequency of the classification label corresponding to keyword a, ∑f a It is the sum of the frequency of occurrence of keywords.
3. The method for classifying and retrieving technical achievements according to claim 2, characterized in that: The steps for obtaining the preliminary classification results are specifically as follows: Based on the keyword field mapping table, access the literature database, extract the keyword list of each document, including database query, data extraction and preliminary data processing, and generate a literature keyword list; Based on the document keyword list, determine the corresponding field label for each keyword, capture the corresponding field label, mark unclassified keywords as unclassified, and generate a mapping result from keywords to fields; Summarize the mapping results of the keywords to fields of each document, and use weighted statistical methods to optimize classification accuracy, through the formula: The preliminary classification results are obtained by calculating the weighted scores of each field, where δ b,k is the indicator function of keyword k, w k is the weight of the keyword, N is the total number of keywords in the document, L b Represents the weighted score of the target domain.
4. The method for classifying and retrieving technical achievements according to claim 3, characterized in that: The steps for obtaining the optimized classification results are specifically as follows: Based on the preliminary classification results, extracting document data from the database, including document identifiers, keyword lists and preliminary field labels, and evaluating the classification of the document data to obtain a classification accuracy evaluation result; According to the classification accuracy evaluation results, by comparing the preliminary domain label of each document with the real annotated domain, the number of instances of successful and failed classification is calculated to obtain the accuracy of the domain label; According to the accuracy of the domain labels, the parameters in the classification algorithm, including weights and thresholds, are adjusted to optimize the overall performance of the classification, using the formula: Generate optimized classification results, where P c represents the prediction accuracy of the optimized domain, f c is the preliminary classification accuracy of the field, s c is an additional parameter adjusted based on domain characteristics, and λ is used to adjust f c and c Regularization coefficient of the contribution.
5. The method for classifying and retrieving technical achievements according to claim 4, characterized in that: The steps for obtaining the new document index architecture are as follows: Based on the optimized classification results, the associated document records are called from the database, including the unique identifier, keywords, original classification and new classification of each document, and the completeness and accuracy of the data are determined using database query commands to generate a document record list; Apply the current index structure update algorithm to the document record list, by formula: I d =hash(K d +C old +C new ) Calculate the new index key value for each document and generate the current index mapping, where I d is the new index, K d is the keyword of the literature, C old and C new They are old classification and new classification respectively; The current index mapping is updated to the database, including replacing the current index in the database and adding a new index, to generate a new document index architecture.
6. The method for classifying and retrieving technical achievements according to claim 5, characterized in that: The steps for obtaining the maintenance and update strategy are specifically as follows: Based on the new document indexing architecture, according to the current document review and index structure, design and implement an index update plan, and generate updated index parameters by comparing and analyzing the differences between the current index and the new requirements; Using the updated index parameters, combined with the real-time user feedback data set, each index item is tested for data integrity and accuracy to generate a difference analysis result; According to the difference analysis results, the index items with insufficient classification effect and user satisfaction are revised, and the formula is applied: The overall deviation of the revised index items is optimized to generate a maintenance update strategy, where p h Represents the original index parameter, q h Represents the revised index parameter, w h represents the weight coefficient, M represents the total number of parameters, and E is the weighted mean square error evaluation value of the revision effect.
7. The method for classifying and retrieving technical achievements according to claim 6, characterized in that: The steps for obtaining the optimized classification retrieval scheme are specifically as follows: Based on the maintenance and update strategy, develop the current literature search interface, integrate search engine technology, conduct keyword and multi-field combined search, and generate a fully functional search interface prototype; Conduct real-time operational simulation of the fully functional search interface prototype, capture preliminary user feedback and query logs, analyze data and identify common query questions and user preferences, adjust search engine configuration parameters, optimize user experience, and generate a preliminary optimized search engine configuration; Based on the above preliminary optimized search engine configuration, the formula is applied: Adjust multiple parameters, optimize query efficiency and accuracy by controlling weight distribution, and generate optimized classification retrieval solutions, where r g represents the real-time response time, t g represents the target response time, w g is the weight set for differentiated query types, Q represents the total error value after optimization, and G represents the total number of query types.
8. A technical achievement classification and retrieval system, characterized in that: According to the method for classifying and retrieving technical achievements according to any one of claims 1 to 7, the system comprises: The keyword analysis module identifies and records the keywords of each document based on the metadata of the documents in the scientific literature database, performs frequency statistics on the keywords, sorts the keywords according to the statistical data, calculates the frequency distribution of the keywords, and generates a keyword frequency distribution table; The keyword mapping module analyzes the relationship between keywords and domain labels based on the keyword frequency distribution table, performs domain label mapping on the keywords, associates the keywords with the target domain labels, and constructs a keyword domain mapping table; The classification and annotation module scans the keywords in the documents through the keyword-domain mapping table, annotates the domain label of each document, uses the label to evaluate the preliminary accuracy of the classification, adjusts and optimizes the classification and annotation parameters, and obtains the optimized classification results; The index update module reconstructs the document index structure in the database according to the optimized classification results, performs index update operations, analyzes user feedback information, adjusts index and classification parameters according to the feedback, and establishes a maintenance update strategy; The retrieval interface development module uses the maintenance and update strategy to develop the current document retrieval interface, integrates the search function of the combination of keywords and document attribute fields, tests and optimizes the search logic, analyzes the query log, adjusts the configuration parameters of the search engine, and generates an optimized classification retrieval plan.
Citation Information
Cited By
Block chain-based cross-border trade electronic document storage and verification method and system
CN120856300A