A semantic retrieval rearrangement method for library knowledge resources

CN122817441APending Publication Date: 2026-09-25SHIJIAZHUANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611124433.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有重排方式虽然能够提高检索效率,但对于敏感场景特征对排序结果的贡献程度、唯一识别属性组合带来的区分风险,以及历史检索日志中体现的稳定暴露风险,仍有进一步精细化处理的空间

Benefits of technology

本发明通过获取馆室知识资源检索请求并获得多个候选馆室知识资源簇,在动态场景参数和阈值控制参数的约束下,区分非敏感重排特征、敏感场景特征和唯一识别属性集合,使语义检索重排过程不仅依据资源内容相关性进行排序,还能够结合访问场景、资源状态和属性组合对候选馆室知识资源簇进行更细致的排序判断,提高馆室知识资源语义检索重排结果与实际馆室业务场景之间的适配性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817441A_ABST
    Figure CN122817441A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of library knowledge resource management and semantic retrieval processing, and particularly relates to a library knowledge resource semantic retrieval rearrangement method, which obtains a library knowledge resource retrieval request and a candidate library knowledge resource cluster, extracts dynamic scene parameters, threshold control parameters and rearrangement features, generates two types of rearrangement results and compares them, determines sensitive dependence, unique identification and leakage risk in combination with a retrieval log, and obtains semantic retrieval rearrangement results according to the regulation and control output. The present application distinguishes between non-sensitive rearrangement features, sensitive scene features and unique identification attributes, forms a rearrangement comparison result, evaluates sensitive dependence, unique identification and sorting stability leakage risk, and regulates the output of the candidate library knowledge resource cluster accordingly, thereby improving the scene adaptability, safety and controllability of semantic retrieval rearrangement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of library knowledge resource management and semantic retrieval processing technology, and more specifically, to a method for semantic retrieval and rearrangement of library knowledge resources. Background Technology

[0002] With the deepening of digitalization in libraries, archives, museums, exhibition halls, laboratories, and general information centers, their knowledge resources have expanded from simple textual materials to include various types of data such as document records, collection descriptions, exhibition materials, restoration records, storage locations, workflow records, personnel responsibility records, and retrieval records. These resources typically exhibit thematic, spatial, business, and temporal connections. Simply relying on keyword matching is insufficient to fully reflect the semantic relationships between resources. Therefore, semantic retrieval and reordering technologies are increasingly being applied to library knowledge resource retrieval scenarios to improve the recall quality and ranking accuracy of candidate library knowledge resource clusters.

[0003] Existing methods for retrieving library knowledge resources typically combine information such as resource title, abstract content, keywords, classification number, tags, citation relationships, and access restrictions to sort and display search results. This approach effectively meets the needs of general information retrieval, collection browsing, business record retrieval, and knowledge association discovery, and can also prevent unauthorized resource content from being directly displayed through access restrictions. However, in library operations, the ranking of search results can itself carry implicit information. For example, a candidate library knowledge resource cluster may receive a higher ranking due to its precise storage location, current storage status, restoration progress, special business processes, or the identity of the responsible personnel. Even if sensitive fields are not directly output, searchers may still infer the specific context and status of the resource by observing its high ranking, changes in position, or stable appearance in multiple similar searches.

[0004] Furthermore, the sensitivity of library knowledge resources does not solely depend on the resource content itself, but is also closely related to dynamic scenario parameters. The visibility and ranking risk of the same resource may differ depending on the access identity, business role, library area, resource access level, query time, and resource circulation status. If the reordering process determines the output order solely based on semantic relevance and static permission conditions, it may be difficult to distinguish between normal relevance enhancements from non-sensitive reordering features and abnormal ranking advantages from sensitive scenario features. Simultaneously, when a candidate library knowledge resource cluster has relatively few combinations of identity, location, status, and time attributes, this cluster may also form a high degree of uniqueness within the same permission range, increasing the likelihood of indirect location or repeated verification.

[0005] Furthermore, the retrieval of knowledge resources within museums is typically continuous and repetitive. Researchers, librarians, or staff may conduct multiple similar searches focusing on the same collection, the same restoration project, the same exhibition theme, or the same spatial area. If a candidate cluster of museum knowledge resources consistently ranks highly in historical re-ranking results, or exhibits low ranking fluctuations under similar search requests, this ranking stability may provide additional inference clues. While existing re-ranking methods can improve retrieval efficiency, there is still room for further refinement in addressing the contribution of sensitive scene features to ranking results, the differentiation risks posed by unique identification attribute combinations, and the stability exposure risks reflected in historical search logs.

[0006] Therefore, this application proposes a semantic retrieval and reordering method for library knowledge resources. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the present invention aims to provide a semantic retrieval and rearrangement method for library knowledge resources.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A semantic retrieval and reordering method for library knowledge resources includes: Step 1: Obtain the library knowledge resource retrieval request and retrieve multiple candidate library knowledge resource clusters corresponding to the library knowledge resource retrieval request from the library knowledge resource index; Step 2: Obtain the dynamic scene parameters and threshold control parameters corresponding to the library knowledge resource retrieval request, and determine the non-sensitive rearrangement features, sensitive scene features and unique identification attribute set based on the dynamic scene parameters and candidate library knowledge resource clusters; Step 3: For each candidate library knowledge resource cluster, generate a first re-ranking score and a first re-ranking result based on non-sensitive re-ranking features and sensitive scene features. Then, generate a second re-ranking score and a second re-ranking result while masking or weakening sensitive scene features and retaining non-sensitive re-ranking features. The re-ranking comparison result is formed by the first re-ranking score, the first re-ranking result, the second re-ranking score, and the second re-ranking result. Step 4: Based on the rearrangement comparison results, the unique identification attribute set, and the pre-recorded retrieval logs, determine the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk of the candidate library knowledge resource clusters. Step 5: Based on the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk, output regulation is applied to the candidate library knowledge resource clusters to obtain the semantic retrieval re-ranking results of library knowledge resources.

[0009] In one embodiment, dynamic scene parameters are used to determine the permission conditions corresponding to the candidate library knowledge resource clusters, and threshold control parameters are used to determine the score difference normalization condition, the ranking change normalization condition, the dependency fusion condition, the similarity retrieval judgment condition, and the perturbation ranking condition.

[0010] In one embodiment, the non-sensitive rearrangement features include at least one of semantic matching features, resource association features, and credibility features; the sensitive scenario features include at least one of spatial location features, resource status features, and business security features; and the unique identification attribute set includes at least one of identity attributes, location attributes, status attributes, and time attributes.

[0011] In one embodiment, the first rearrangement score and the second rearrangement score use the same rearrangement score calculation structure so that the difference between the first rearrangement score and the second rearrangement score characterizes the degree of influence of sensitive scene features on the ranking result.

[0012] In one embodiment, the method for determining the ranking dependency sensitivity feature includes: determining the score sensitivity offset value and the position sensitivity offset value of the candidate library knowledge resource cluster based on the re-ranking comparison result, and fusing the score sensitivity offset value and the position sensitivity offset value according to the dependency fusion condition to obtain the ranking dependency sensitivity feature.

[0013] In one embodiment, the determination of the score-sensitive offset value and the position-sensitive offset value includes: extracting the first rearrangement score and the second rearrangement score from the rearrangement comparison results, as well as the first and second positions of the candidate library knowledge resource clusters in the first and second rearrangement results; obtaining the score-sensitive offset value according to the score difference normalization condition based on the difference between the first rearrangement score and the second rearrangement score; and obtaining the position-sensitive offset value according to the position change normalization condition based on the difference between the first and second positions.

[0014] In one embodiment, the unique identification degree of a resource cluster is calculated as follows: Based on the unique identification attribute set, candidate library knowledge resource clusters with the same permissions that match the candidate library knowledge resource clusters in the corresponding dimension of the unique identification attribute set and meet the permission conditions determined by the dynamic scene parameters are identified in the library knowledge resource index, and the number of candidate library knowledge resource clusters with the same permissions is counted; based on the number of candidate library knowledge resource clusters with the same permissions and the unique identification attribute set, the resource cluster unique identification degree of the candidate library knowledge resource clusters is calculated.

[0015] In one embodiment, based on the library knowledge resource retrieval request and the pre-recorded retrieval log, historical retrieval requests that meet the similarity retrieval judgment conditions determined by the threshold control parameters and the historical re-ranking results corresponding to the historical retrieval requests are extracted; the historical ranking fluctuation value and historical front frequency of the candidate library knowledge resource cluster are determined based on the historical re-ranking results, and the ranking stability leakage risk of the candidate library knowledge resource cluster is calculated based on the ranking dependency sensitivity feature degree, resource cluster uniqueness, historical ranking fluctuation value and historical front frequency.

[0016] In one embodiment, step five includes: generating the regulation results of candidate library knowledge resource clusters based on the ranking dependency sensitivity feature degree, resource cluster uniqueness identification degree and ranking stability leakage risk, and determining the output order of candidate library knowledge resource clusters based on the regulation results to obtain the library knowledge resource semantic retrieval reordering results.

[0017] In one embodiment, the generation of the control result includes: determining the reordering control intensity of candidate library knowledge resource clusters based on the ranking dependency sensitivity feature, resource cluster uniqueness, and ranking stability leakage risk, and generating the final reordering score based on the reordering control intensity; determining whether to perform perturbation ranking on candidate library knowledge resource clusters based on the perturbation ranking conditions determined by the threshold control parameters, and generating a set of safe and equivalent candidate library knowledge resource clusters and a constrained perturbation ranking result when it is determined to perform perturbation ranking.

[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention obtains multiple candidate library knowledge resource clusters by acquiring library knowledge resource retrieval requests and, under the constraints of dynamic scene parameters and threshold control parameters, distinguishes between non-sensitive reordering features, sensitive scene features, and a unique set of identification attributes. This enables the semantic retrieval reordering process to not only sort based on the relevance of resource content, but also to make more detailed sorting judgments on candidate library knowledge resource clusters by combining access scenarios, resource status, and attribute combinations, thereby improving the adaptability of the semantic retrieval reordering results of library knowledge resources to actual library business scenarios. For each candidate library knowledge resource cluster, a first re-ranking score and a first re-ranking result containing sensitive scene features are generated, as well as a second re-ranking score and a second re-ranking result under the condition of masking or weakening sensitive scene features while retaining non-sensitive re-ranking features. The re-ranking comparison results are formed in this way, which can more accurately determine the degree of influence of sensitive scene features on the ranking score and ranking position of candidate library knowledge resource clusters, and reduce the risk of exposure of implicit information caused by the excessive participation of sensitive scene features in the ranking. By combining the reordering comparison results, the unique identification attribute set, and the pre-recorded retrieval logs, the ranking dependence sensitivity, resource cluster uniqueness, and ranking stability leakage risk of candidate library knowledge resource clusters are determined. Based on this, the output of candidate library knowledge resource clusters can be controlled. This can suppress the long-term or stable exposure of resource clusters with high uniqueness in similar searches while maintaining the relevance of library knowledge resource retrieval and permission adaptability, thereby improving the security and controllability of the semantic retrieval reordering results of library knowledge resources. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall process of the semantic retrieval and rearrangement method for library knowledge resources according to the present invention; Figure 2 This is a schematic diagram illustrating the process of generating the rearrangement comparison results of the present invention; Figure 3 This is a schematic diagram of the output control process for the candidate library knowledge resource cluster of the present invention. Detailed Implementation

[0020] Reference Figure 1 A semantic retrieval and reordering method for library knowledge resources, comprising: Step 1: Obtain the library's knowledge resource retrieval request and retrieve multiple candidate library knowledge resource clusters corresponding to the retrieval request from the library's knowledge resource index; clarify the retrieval entry point and candidate object scope of the library's knowledge resource semantic retrieval reordering method. The sources of library knowledge resource retrieval requests include librarian terminals, reader terminals, library business systems, or retrieval service interfaces, and their content includes one or more of the following: natural language queries, keyword combinations, collection identifiers, spatial ranges, time ranges, or business scenario information. The library knowledge resource index records the index relationships between library knowledge resource clusters and semantic content, resource attributes, associations, and status information. Candidate library knowledge resource clusters are not single, isolated resource entries, but rather collections of resources that have aggregation relationships with subject content, collection objects, library spaces, operational records, or knowledge nodes. This step completes the definition of the search request, index scope, and candidate set, ensuring that the reordering process has a clearly defined object.

[0021] The system receives library knowledge resource retrieval requests and parses the query statements, resource limitations, library business scenario information, access identity information, and time range information in the retrieval requests to obtain semantic retrieval elements relevant to the current retrieval. Based on these semantic retrieval elements, it performs matching searches in the library knowledge resource index. The library knowledge resource index pre-records the correspondence between library collections, exhibition materials, storage archives, equipment information, spatial location records, business flow records, and their semantic tags, and clusters the library knowledge resources according to subject content, resource source, business relationship, and library spatial relationship. During the matching process, the semantic content in the library knowledge resource retrieval request is comprehensively matched with the resource semantic tags, attribute fields, and related nodes in the library knowledge resource index to select resource sets that have semantic, business, or spatial relationships with the library knowledge resource retrieval request, forming multiple candidate library knowledge resource clusters. For example, when a librarian enters a search request for library knowledge resources related to the restoration records of precious ancient books and the location of the storage room, the library knowledge resource index not only retrieves document resources containing the restoration records of ancient books, but also retrieves the storage room where the ancient book is located, the restoration process records, the transfer records of the responsible librarians, and the records of the same batch of library collection resources. It also aggregates resources with related relationships into different candidate library knowledge resource clusters, which serve as candidate objects for subsequent semantic retrieval and ranking processing.

[0022] Step 2: Obtain the dynamic scene parameters and threshold control parameters corresponding to the library knowledge resource retrieval request, and determine the non-sensitive rearrangement features, sensitive scene features, and unique identification attribute set based on the dynamic scene parameters and candidate library knowledge resource clusters; introduce dynamic scene parameters and threshold control parameters on the basis of candidate library knowledge resource clusters to clarify the feature categories and identification attribute ranges involved in the rearrangement process; Dynamic scenario parameters reflect differences in scenarios such as the current business status of the library / room, access permissions, spatial environment, resource flow status, or query initiator conditions. Threshold control parameters provide the judgment boundaries for score differences, ranking changes, similarity searches, and perturbation ranking. Non-sensitive re-ranking features retain the general relevance between library / room knowledge resources and library / room knowledge resource retrieval requests. Sensitive scenario features correspond to scenario factors that may affect ranking and are related to spatial location, resource status, or business security. The unique identification attribute set records the attribute dimensions that can distinguish candidate library / room knowledge resource clusters separately. This division places relevance evaluation, scenario-sensitive factors, and unique identification attributes within different data dimensions for processing, avoiding mixed calculations.

[0023] The library's knowledge resource retrieval request is denoted as Dynamic scene parameters are denoted as The threshold control parameter is denoted as Dynamic scene parameters The set of candidate library knowledge resource clusters filtered according to the corresponding permission conditions is denoted as

[0024] in, The number of knowledge resource clusters in candidate libraries. For the first A cluster of knowledge resources for candidate museum rooms. Candidate Library Room Knowledge Resource Cluster The insensitive rearrangement eigenvector is denoted as

[0025] The feature vector of a sensitive scene is denoted as

[0026] The set of unique identification attributes is denoted as

[0027] in, This refers to the number of non-sensitive rearrangement feature dimensions. The number of feature dimensions for sensitive scenes, To uniquely identify the number of attribute dimensions, For the first Normalized values ​​of insensitive rearrangement features, , For the first Normalized values ​​of features in sensitive scenarios, , For the first The attribute value that uniquely identifies an attribute. ,and , Threshold control parameters Including score difference normalization parameters Normalized parameters of position variation Similarity retrieval criteria Historical ranking fluctuation normalization parameter Safety equivalence score difference parameter Non-sensitive contribution difference parameters and the first digit truncation value ,in , , , , , , Values ​​are positive integers. The values ​​of non-sensitive rearrangement features and sensitive scenario features are obtained from the semantic content, association relationships, credibility information, and scenario metadata in the library's knowledge resource index. Semantic matching features are obtained based on the semantic similarity between the library's knowledge resource retrieval request and the candidate library's knowledge resource cluster; resource association features are obtained based on at least one of the following association relationships: topic association, spatial association, business association, or temporal association; credibility features are obtained based on at least one of the following: resource source, verification status, record completeness, or update time; spatial location features, resource status features, and business security features are obtained based on the corresponding location field, status field, and business security field in the library's knowledge resource index, and are normalized according to a preset value range. Weights Shielding or weakening coefficient and dependency fusion coefficient The settings are pre-configured by the library management system, configured by the library security administrator, or determined based on historical search samples.

[0028] The system acquires dynamic scene parameters and threshold control parameters corresponding to the library's knowledge resource retrieval request. It reads the access identity, business role, library area, resource access level, current business process, query time, and resource flow status from the dynamic scene parameters, and determines the permission conditions for each candidate library knowledge resource cluster based on these parameters. Simultaneously, it reads the threshold control parameters, which correspond to score difference normalization conditions, ranking change normalization conditions, dependency fusion conditions, similarity retrieval judgment conditions, and perturbation ranking conditions. These provide a unified boundary for calculating rearrangement differences, judging ranking changes, sensitive dependency fusion, historical retrieval matching, and perturbation ranking judgment for candidate library knowledge resource clusters. After determining the permission conditions, non-sensitive rearrangement features, sensitive scenario features, and a set of unique identification attributes are extracted based on the resource content, index fields, business records, and spatial relationships of the candidate library knowledge resource clusters. Non-sensitive rearrangement features include at least one of semantic matching features, resource association features, and credibility features, such as the degree of matching between the library's knowledge resource retrieval request and the resource's subject terms, abstract content, collection classification, citation relationships, and reliable source records. Sensitive scenario features include at least one of spatial location features, resource status features, and business security features, such as the location of valuable documents in storage, whether the resource is in a borrowed, repaired, sealed, or inventory state, and whether the business matter to which the resource belongs involves security management requirements. The set of unique identification attributes includes at least one of identity attributes, location attributes, status attributes, and time attributes, such as the responsible librarian's identity, storage room number, resource status marker, and last access time. For example, when a library knowledge resource retrieval request involves restoration data of a batch of undisclosed cultural relics, and the dynamic scene parameters show that the query initiator only has ordinary cataloging permissions, information related to the precise storage location, current security status, and specific responsible personnel of the batch of cultural relics is classified into sensitive scene features or unique identification attribute sets, while the cultural relic's age, subject classification, restoration process summary, and public collection description are classified into non-sensitive rearrangement features.

[0029] Reference Figure 2 Step 3: For each candidate library knowledge resource cluster, generate a first re-ranking score and a first re-ranking result based on non-sensitive re-ranking features and sensitive scene features. Then, generate a second re-ranking score and a second re-ranking result while masking or weakening sensitive scene features and retaining non-sensitive re-ranking features. The re-ranking comparison result is formed by the first re-ranking score, the first re-ranking result, the second re-ranking score, and the second re-ranking result. The first re-ranking score and the first re-ranking result reflect the ranking status when non-sensitive re-ranking features and sensitive scene features are involved together. The second re-ranking score and the second re-ranking result are formed under the condition of masking or weakening sensitive scene features while retaining non-sensitive re-ranking features, thus preserving the semantic relevance while reducing the direct impact of sensitive scene features on the score and ranking. The re-ranking comparison result formed by the first re-ranking score, the first re-ranking result, the second re-ranking score, and the second re-ranking result records the score and ranking differences of the same candidate library knowledge resource cluster under the two calculation conditions, serving as the basis for analyzing the sources of ranking changes.

[0030] For each candidate library knowledge resource cluster, non-sensitive rearrangement features and sensitive scenario features are input into the same rearrangement score calculation structure. The first rearrangement score of the corresponding candidate library knowledge resource cluster is obtained by comprehensively calculating the semantic matching degree, resource association strength, credibility weight, spatial location influence, resource status influence and business security influence. The first rearrangement result is formed based on the first rearrangement score of each candidate library knowledge resource cluster. While maintaining the same re-ranking score calculation structure, the same candidate library knowledge resource cluster range, and the same non-sensitive re-ranking feature values, sensitive scene features are masked or weakened. Masking includes removing spatial location features, resource status features, or business security features from the score calculation, while weakening includes reducing the weight of sensitive scene features, expanding the value range of sensitive scene features, or replacing precise attributes with hierarchical attributes. Under these conditions, the second re-ranking score is calculated, and the second re-ranking result is formed based on the second re-ranking score. The first re-ranking score, the first re-ranking result, the second re-ranking score, and the second re-ranking result are recorded as the re-ranking comparison results. The first and second ranking scores use the same linear normalization calculation structure. Candidate library knowledge resource clusters. The score for the first rearrangement is denoted as Its calculation formula is

[0031] Candidate Library Room Knowledge Resource Cluster The score for the second rearrangement is denoted as Its calculation formula is

[0032] in, For the first The weights of insensitive rearrangement features, For the first Weights of features in sensitive scenarios, For the first The masking or weakening coefficient of a sensitive scene feature. For the first The values ​​of sensitive scene features after being masked or weakened. , , , ,and

[0033] When the When a sensitive scene feature is masked When the first When the features of a sensitive scene are weakened , The ranking is obtained through hierarchical mapping of precise attributes; precise storage location is mapped to the library or storage area level attribute, precise resource status is mapped to the open, processing, and restricted access level attribute, and precise business security level is mapped to the public level attribute. The calculation of the second re-ranking score does not change the set of candidate library knowledge resource clusters. The boundaries of permissions.

[0034] For example, when a library's knowledge resource retrieval request is for a batch of restored materials, a candidate library knowledge resource cluster ranks highly in the first ranking result because it includes the precise storage location, current restoration status, and business security level. However, when the precise storage location is masked, the current restoration status is weakened to "under general processing," and the business security level is only calculated based on the public classification, the candidate library knowledge resource cluster's second ranking score decreases, and its ranking changes. Since the first and second ranking scores use the same ranking score calculation structure, this score difference and ranking change can reflect the degree of influence of sensitive scene characteristics on the ranking results.

[0035] Step four: Based on the re-ranking comparison results, the unique identification attribute set, and the pre-recorded retrieval logs, determine the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk of the candidate library knowledge resource clusters. This step combines the re-ranking comparison results, the unique identification attribute set, and the pre-recorded retrieval logs for examination to obtain the sensitivity dependency, distinguishability, and historical stability risk of the candidate library knowledge resource clusters during the ranking process. Ranking dependency sensitivity reflects the degree to which the ranking results of candidate library knowledge resource clusters depend on sensitive scenario features. Resource cluster uniqueness reflects the degree to which candidate library knowledge resource clusters are individually located or distinguished under the corresponding dimension of the unique identification attribute set. Ranking stability leakage risk reflects the possibility of information leakage when candidate library knowledge resource clusters repeatedly occupy a high position or experience abnormal ranking fluctuations in similar retrieval scenarios. Pre-recorded retrieval logs provide historical retrieval requests, historical re-ranking results, and corresponding ranking changes; risk assessment is not limited to a single retrieval result. This step incorporates sensitivity dependency, unique identification, and historical stability into the same risk analysis process.

[0036] The ranking dependency sensitivity of candidate library knowledge resource clusters is determined based on the rearrangement comparison results. Specifically, the first rearrangement score, second rearrangement score, first ranking, and second ranking of the same candidate library knowledge resource cluster are read from the rearrangement comparison results. The difference between the first rearrangement score and the second rearrangement score is calculated, and the difference is converted into a score-sensitive offset value according to the score difference normalization condition determined by the threshold control parameter. At the same time, the difference between the first ranking and the second ranking is calculated, and the difference is converted into a ranking-sensitive offset value according to the ranking change normalization condition determined by the threshold control parameter. With both the score-sensitive offset value and the ranking-sensitive offset value obtained, they are weighted and fused according to the dependency fusion condition to obtain the ranking dependency sensitivity of the candidate library knowledge resource cluster. Candidate Library Room Knowledge Resource Cluster The position in the first rearrangement is denoted as The position in the second rearrangement is denoted as The smaller the digit value, the higher the corresponding sorting position. Candidate Library Knowledge Resource Clusters The score sensitivity offset value is denoted as Its calculation formula is

[0037] Candidate Library Room Knowledge Resource Cluster The position-sensitive offset value is denoted as Its calculation formula is

[0038] Candidate Library Room Knowledge Resource Cluster The ranking dependency sensitivity feature is denoted as Its calculation formula is

[0039] in, The parameter is used to normalize the score difference. The parameter is used to normalize the position variation. The dependency fusion coefficient, . The higher the value, the better the candidate library knowledge resource cluster. The higher the ranking result depends on the features of sensitive scenes.

[0040] The unique identification degree and sorting stability leakage risk of candidate library knowledge resource clusters are determined based on the unique identification attribute set and pre-recorded retrieval logs. Specifically, based on the identity attribute, location attribute, status attribute, and time attribute in the unique identification attribute set, candidate library knowledge resource clusters with the same permissions that match the candidate library knowledge resource clusters in the same attribute dimension and meet the permission conditions determined by dynamic scenario parameters are searched in the library knowledge resource index, and the number of candidate library knowledge resource clusters with the same permissions is counted. The fewer the number of candidate library knowledge resource clusters with the same permissions, the higher the degree to which the candidate library knowledge resource clusters are distinguished individually in the corresponding attribute dimension, and the unique identification degree of the resource cluster is increased accordingly. Candidate Library Room Knowledge Resource Cluster The A unique identification attribute is , In dynamic scene parameters Under the same permission conditions, with In the The number of candidate library knowledge resource clusters that match on a unique identification attribute is denoted as . ;and In the unique identification attribute set The number of candidate library knowledge resource clusters that match across all attribute dimensions is denoted as The statistics include candidate library knowledge resource clusters. itself, therefore , Candidate Library Room Knowledge Resource Cluster The unique identifier of a resource cluster is denoted as Its calculation formula is

[0041] in, Weights are assigned to the number of attributes across all dimensions. For the first The weight of the number of matches for each unique identification attribute dimension , ,and

[0042] The fewer matches within the same permission range, the better. or The higher the value, the better the candidate library knowledge resource cluster. The higher the uniqueness of the resource cluster, the better.

[0043] Meanwhile, based on the library's knowledge resource retrieval requests, historical retrieval requests that meet the similarity retrieval criteria are extracted from the pre-recorded retrieval logs. Historical re-ranking results corresponding to these historical retrieval requests are obtained, and the historical ranking fluctuation value and historical top frequency of candidate library knowledge resource clusters in the historical re-ranking results are statistically analyzed. The risk of ranking stability leakage is calculated by combining ranking dependency sensitivity features, resource cluster uniqueness, historical ranking fluctuation value, and historical top frequency. For example, in multiple similar searches related to the restoration of precious ancient books, if a candidate library knowledge resource cluster consistently ranks highly due to a combination of specific storage location, current storage status, and responsible librarian's identity, and only a small number of resource clusters with the same attribute combination exist within the same access scope, then the resource cluster uniqueness and ranking stability leakage risk of this candidate library knowledge resource cluster are judged to be high.

[0044] In this embodiment, the pre-recorded set of retrieval logs is denoted as

[0045] in, For the first 1 historical search request, For the first Historical reordering results corresponding to each historical search request. For the first The search time of each historical search request. For the first Dynamic scene parameters corresponding to each historical search request. This represents the number of log entries. The current library's knowledge resource retrieval requests. The semantic vector is denoted as Historical search requests The semantic vector is denoted as Both are obtained from the same semantic vector generation model; dynamic scene parameters The corresponding permission conditions are denoted as The set of historical search requests that meet the similarity retrieval criteria is denoted as .

[0046] in, For similarity retrieval determination parameters, Candidate Library Room Knowledge Resource Cluster Appearing in the set The number of times corresponding to the historical rearrangement results is denoted as Before entering the historical rearrangement results The number of digits is denoted as Candidate Library Room Knowledge Resource Cluster Historical pre-position frequency is recorded as Its calculation formula is

[0047] Candidate Library Room Knowledge Resource Cluster In the set The standard deviation of the position in the corresponding historical rearrangement results is denoted as Candidate Library Room Knowledge Resource Cluster Historical ranking fluctuation value is denoted as Its calculation formula is

[0048] Candidate Library Room Knowledge Resource Cluster The risk of leakage of sorting stability is denoted as Its calculation formula is

[0049] in, The parameter is used to normalize the historical ranking fluctuations. The truncation value is the preceding position. For risk fusion weights, and , It is a positive integer. , , , , . The smaller the size, the more stable its historical ranking. The higher the value, the greater the risk of sorting stability leakage. The higher.

[0050] Reference Figure 3 Step five involves adjusting the output of candidate library knowledge resource clusters based on ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk. This yields a semantic retrieval re-ranking result for library knowledge resources. The ranking performance of the candidate library knowledge resource clusters is then adjusted at the output stage of this re-ranking process. This output adjustment does not alter the correspondence between library knowledge resource retrieval requests and candidate library knowledge resource clusters; rather, it processes the scores, order, or perturbation range of the candidate library knowledge resource clusters at the ranking output stage. When ranking depends on sensitive feature degree, resource cluster uniqueness, or ranking stability leakage risk, indicating that candidate library knowledge resource clusters pose a high risk, output adjustment reduces the impact of sensitive scene features on the ranking, minimizing the stable exposure of highly unique resource clusters in the search results. The semantic retrieval re-ranking results for library knowledge resources maintain the relevance of the library knowledge resource retrieval request while imposing constraints on candidate library knowledge resource clusters involving sensitive scene features, unique identification attribute sets, and historical ranking stability.

[0051] The regulation results for candidate library knowledge resource clusters are generated based on the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk. Specifically, for the same candidate library knowledge resource cluster, its ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk are read, and the re-ranking regulation intensity is determined based on these factors. When the ranking dependency sensitivity is high, it indicates that the ranking advantage of the candidate library knowledge resource cluster mainly comes from sensitive scene features, and the re-ranking regulation intensity is increased accordingly. When the uniqueness of a resource cluster is high, it indicates that the candidate library knowledge resource cluster is easily distinguishable by its identity, location, status, or time attributes, and the intensity of the reordering control is further increased. When the risk of leakage of ranking stability is high, it indicates that the candidate library knowledge resource cluster has maintained a leading position or exhibits an abnormally stable ranking exposure state in similar searches for a long time, and the control results restrict the degree of leading position of the candidate library knowledge resource cluster.

[0052] Candidate Library Room Knowledge Resource Cluster The contribution value of the insensitive rearrangement feature is denoted as Its calculation formula is

[0053] Candidate Library Room Knowledge Resource Cluster The rearrangement regulation intensity is denoted as Its calculation formula is

[0054] in, To adjust the fusion weight, , , ,and Candidate Library Room Knowledge Resource Cluster The final rearrangement score is recorded as Its calculation formula is

[0055] Candidate Library Room Knowledge Resource Cluster The regulation result is recorded as ,and

[0056] when hour, ;when hour, The basic output order is as follows: Sort in descending order; The same, according to Sort in descending order; and If they are all the same, they are arranged in ascending order according to the fixed number of the knowledge resource cluster of the candidate library.

[0057] Based on the perturbation ranking conditions determined by the threshold control parameters, it is determined whether to perturb the ranking of candidate library knowledge resource clusters. When the difference in the final re-ranking scores among multiple candidate library knowledge resource clusters is within the range limited by the perturbation ranking conditions, and their semantic matching features, resource association features, and credibility features meet the security equivalence requirements, these candidate library knowledge resource clusters are included in the set of safe and equivalent candidate library knowledge resource clusters. Within the set of safe and equivalent candidate library knowledge resource clusters, the local output order is adjusted according to the constrained perturbation ranking rules. The adjustment range does not cross the permission conditions, does not increase the exposure level of high-risk resource clusters, and does not destroy the basic relevance between the library knowledge resource retrieval request and the candidate library knowledge resource clusters. In this embodiment, candidate library knowledge resource clusters in the basic output order are used. Based on the ensemble, candidate library knowledge resource clusters The final rearrangement score is recorded as The contribution value of insensitive rearrangement features is denoted as Dynamic scene parameters Next candidate library room knowledge resource cluster The permission conditions are denoted as .and The corresponding set of security-equivalent candidate library knowledge resource clusters is denoted as Its definition

[0058] in, For the safety equivalence score difference parameter, For non-sensitive contribution difference parameters, , .when The perturbation sorting condition is met when the number of knowledge resource clusters in the candidate library is no less than two; when When the number of knowledge resource clusters in the candidate library is less than two, the set will not undergo local reordering. The candidate library knowledge resource clusters within the set are sorted in the following order: ascending order of rearrangement intensity, descending order of contribution value of non-sensitive rearrangement features, descending order of final rearrangement score, and ascending order of fixed numbers of candidate library knowledge resource clusters. The fixed numbers of candidate library knowledge resource clusters are unique numbers pre-recorded in the library knowledge resource index. Candidate library knowledge resource clusters outside the set of safe and equivalent candidate library knowledge resource clusters maintain the basic output order; Local reordering within an instance does not cross permission conditions and does not change permissions. Candidate library knowledge resource clusters that do not meet the permission conditions will not be included in the output results.

[0059] For example, when a librarian searches for records of the restoration of precious ancient books, three candidate library knowledge resource clusters are all highly related to the restoration process, restoration batch, and collection period in terms of semantic content. One candidate library knowledge resource cluster has been ranked highly for a long time because it contains the precise storage location and current sealing time. After calculation, its ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk are all high. Therefore, this candidate library knowledge resource cluster receives a greater re-ranking control intensity. Finally, the re-ranking score is constrained within the set of safe and equivalent candidate library knowledge resource clusters, and the ranking result is output after constrained perturbation, resulting in the semantic retrieval re-ranking result of library knowledge resources.

[0060] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A semantic retrieval and reordering method for library knowledge resources, characterized in that, include: Step 1: Obtain the library knowledge resource retrieval request and retrieve multiple candidate library knowledge resource clusters corresponding to the library knowledge resource retrieval request from the library knowledge resource index; Step 2: Obtain the dynamic scene parameters and threshold control parameters corresponding to the library knowledge resource retrieval request, and determine the non-sensitive rearrangement features, sensitive scene features and unique identification attribute set based on the dynamic scene parameters and candidate library knowledge resource clusters; Step 3: For each candidate library knowledge resource cluster, generate a first re-ranking score and a first re-ranking result based on non-sensitive re-ranking features and sensitive scene features. Then, generate a second re-ranking score and a second re-ranking result while masking or weakening sensitive scene features and retaining non-sensitive re-ranking features. The re-ranking comparison result is formed by the first re-ranking score, the first re-ranking result, the second re-ranking score, and the second re-ranking result. Step 4: Based on the rearrangement comparison results, the unique identification attribute set, and the pre-recorded retrieval logs, determine the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk of the candidate library knowledge resource clusters. Step 5: Based on the ranking dependency sensitivity, resource cluster uniqueness, and ranking stability leakage risk, output regulation is applied to the candidate library knowledge resource clusters to obtain the semantic retrieval re-ranking results of library knowledge resources.

2. The semantic retrieval and rearrangement method for library knowledge resources according to claim 1, characterized in that, Dynamic scene parameters are used to determine the permission conditions corresponding to the candidate library knowledge resource clusters, while threshold control parameters are used to determine the score difference normalization condition, the ranking change normalization condition, the dependency fusion condition, the similarity retrieval judgment condition, and the perturbation ranking condition.

3. The semantic retrieval and reordering method for library knowledge resources according to claim 2, characterized in that, Non-sensitive rearrangement features include at least one of semantic matching features, resource association features, and credibility features; sensitive scenario features include at least one of spatial location features, resource status features, and business security features; the unique identification attribute set includes at least one of identity attributes, location attributes, status attributes, and time attributes.

4. The semantic retrieval and reordering method for library knowledge resources according to claim 3, characterized in that, The first and second rearrangement scores use the same rearrangement score calculation structure so that the difference between the first and second rearrangement scores represents the degree of influence of sensitive scene features on the ranking results.

5. The semantic retrieval and reordering method for library knowledge resources according to claim 4, characterized in that, The method for determining the ranking dependency sensitivity feature includes: determining the score sensitivity offset and position sensitivity offset of the candidate library knowledge resource clusters based on the re-ranking comparison results, and fusing the score sensitivity offset and position sensitivity offset according to the dependency fusion condition to obtain the ranking dependency sensitivity feature.

6. The semantic retrieval and reordering method for library knowledge resources according to claim 5, characterized in that, The methods for determining the score-sensitive offset and the position-sensitive offset include: extracting the first and second re-rank scores from the re-ranking comparison results, as well as the first and second positions of the candidate library knowledge resource clusters in the first and second re-ranking results; obtaining the score-sensitive offset based on the difference between the first and second re-rank scores according to the score difference normalization condition; and obtaining the position-sensitive offset based on the difference between the first and second positions according to the position change normalization condition.

7. The semantic retrieval and reordering method for library knowledge resources according to claim 1, characterized in that, The calculation method for the uniqueness of resource clusters is as follows: Based on the unique identification attribute set, identify the candidate library knowledge resource clusters with the same permissions in the library knowledge resource index that match the candidate library knowledge resource clusters in the corresponding dimension of the unique identification attribute set and meet the permission conditions determined by the dynamic scenario parameters, and count the number of candidate library knowledge resource clusters with the same permissions. Based on the number of candidate library knowledge resource clusters with the same access level and the set of unique identification attributes, calculate the resource cluster uniqueness of the candidate library knowledge resource clusters.

8. The semantic retrieval and reordering method for library knowledge resources according to claim 7, characterized in that, Based on the library's knowledge resource retrieval requests and pre-recorded retrieval logs, historical retrieval requests that meet the similarity criteria determined by the threshold control parameters and the corresponding historical re-ranking results are extracted. Based on the historical re-ranking results, the historical ranking fluctuation value and historical front frequency of candidate library knowledge resource clusters are determined, and the ranking stability leakage risk of candidate library knowledge resource clusters is calculated based on the ranking dependency sensitivity feature, resource cluster uniqueness, historical ranking fluctuation value, and historical front frequency.

9. A semantic retrieval and rearrangement method for library knowledge resources according to any one of claims 1-8, characterized in that, Step five includes: generating adjustment results for candidate library knowledge resource clusters based on ranking dependency sensitivity features, resource cluster uniqueness, and ranking stability leakage risk; determining the output order of candidate library knowledge resource clusters based on the adjustment results; and obtaining the semantic retrieval reordering results of library knowledge resources.

10. A semantic retrieval and rearrangement method for library knowledge resources according to claim 9, characterized in that, The generation of the control results includes: determining the reordering control intensity of candidate library knowledge resource clusters based on the ranking dependency sensitivity feature, resource cluster uniqueness, and ranking stability leakage risk, and generating the final reordering score based on the reordering control intensity; determining whether to perform perturbation ranking on candidate library knowledge resource clusters based on the perturbation ranking conditions determined by the threshold control parameters, and generating a safe and equivalent candidate library knowledge resource cluster set and a constrained perturbation ranking result when it is determined to perform perturbation ranking.