Personalized document retrieval method and device, equipment and storage medium

The individualized document retrieval method addresses inefficiencies in traditional retrieval by generating personalized queries with time-sensitive user interests, improving efficiency and accuracy through a lightweight scoring model.

CN120316249APending Publication Date: 2025-07-15SHANDONG UNIV +2

Patent Information

Application Number
CN202510783149.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to user personalized needs in data retrieval, resulting in low retrieval efficiency and accuracy.

Method used

By receiving user's document search requests, determine user query information and job information, generate personalized user query information, combine time-sensitive user dynamic interest information, search relevant documents from a multi-dimensional database, and use the target scoring model to calculate the correlation score, and select the document that is most suitable for user needs.

Benefits of technology

It improves the efficiency and accuracy of data retrieval, reduces the consumption of computing resources, and provides a smoother user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316249A_ABST
    Figure CN120316249A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized document retrieval method, device and equipment and a storage medium, and relates to the technical field of data processing, and the personalized document retrieval method comprises the following steps: generating personalized user query information according to user query information, user post information and time-sensitive user dynamic interest information; calculating a correlation score between the personalized user query information and each related document based on the target scoring model; and determining a matched personalized document according to the correlation score. Through the above mode, fine fluctuation of user interest is captured according to a mode of generating time-sensitive user dynamic interest information, personalized user query information is generated in combination with query information and post information, and then the lightweight target scoring model is adopted to calculate the correlation score, so that the retrieval efficiency is improved while the retrieval effect is maintained. According to the method, the consumption of computing resources is remarkably reduced, and then the personalized document is determined according to the correlation score, so that the document retrieval efficiency and accuracy can be effectively improved, and smoother interaction experience is provided for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a personalized document retrieval method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of information technology, the sharp growth of data volume has brought unprecedented challenges to individuals and enterprises. The data contains rich information and potential value. Especially enterprise data often has obvious domain characteristics and sensitivity. Traditional data retrieval methods mainly rely on publicly available data for modeling, and cannot process professional concepts and knowledge in domain data. At the same time, due to the lack of labeled data, it is unable to meet the personalized needs of users, and it is difficult to adapt to the dynamic changes of user needs. If directly retrieving, the efficiency and accuracy of retrieved documents will be relatively low.

[0003] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of the present application is to provide a personalized document retrieval method, apparatus, device, and storage medium, aiming to solve the technical problem of relatively low efficiency and accuracy of retrieving documents in the prior art.

[0005] To achieve the above purpose, the present application proposes a personalized document retrieval method, and the method includes: When receiving a document retrieval request from a user, determining user query information and user position information according to the document retrieval request; Generating personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information; Retrieving relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculating a relevance score between the personalized user query information and each of the relevant documents based on a target scoring model; Determining a personalized document that matches the document retrieval request according to the relevance score.

[0006] In one embodiment, the step of generating personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information includes: Obtaining historical behavior information of the user from a target historical record storage database; Performing keyword extraction on the historical behavior information based on a professional concept dictionary; Generating time-sensitive user dynamic interest information according to historical behavior keywords; Generate personalized user query information based on the user query information, the user position information, and the time-sensitive user dynamic interest information.

[0007] In one embodiment, the step of generating time-sensitive user dynamic interest information based on historical behavior keywords includes: Set an initial weight for the historical behavior keywords; Obtain the number of occurrences of the historical behavior keywords; Adjust the initial weight according to the number of occurrences and the target decay factor to obtain the target weight; Sort the historical behavior keywords according to the target weight, and generate time-sensitive user dynamic interest information according to the sorting result.

[0008] In one embodiment, before the step of calculating the relevance score between the personalized user query information and each of the relevant documents based on the target scoring model, the method further includes: Obtain positive sample documents related to the user query information and negative sample documents unrelated to the user query information; Based on a pre-trained lightweight re-ranking model, train the current scoring model according to the target scoring function, the positive sample documents, and the negative sample documents; Calculate the loss value of the current scoring model according to the contrast loss function and the preset temperature parameter; Perform validation training on the current scoring model according to the preset validation samples, and determine the target scoring model when the loss value converges to a preset value.

[0009] In one embodiment, the step of determining the target scoring model when the loss value converges to a preset value includes: When the loss value converges to a preset value, obtain the document retrieval result output by the target large model; Perform semantic analysis on each document in the document retrieval result, and perform relevance analysis on each document after semantic analysis to generate a personalized document sorted list; Count the number of documents in the personalized document sorted list, and construct a target document pair according to the number; Fine-tune the current scoring model after validation training according to the target document pair to obtain the target scoring model.

[0010] In one embodiment, the step of determining the user query information and the user position information according to the document retrieval request when receiving the user's document retrieval request includes: When receiving the user's document retrieval request, parse the document retrieval request; Obtain the retrieval request header information according to the parsing result; When the verification of the retrieval request header information passes, obtain the retrieval request content according to the parsing result, and determine the user query information according to the retrieval request content; Obtain the login information of the device that initiated the document retrieval request according to the parsing result; Determine the user position information according to the login information.

[0011] In one embodiment, the step of determining the personalized document that matches the document retrieval request according to the relevance score includes: Generate a candidate document list according to each relevant document corresponding to the relevance score; Intelligently sort the candidate document list according to the relevance score; Respond to the document selection instruction triggered by the user on the document display page, and select the personalized document that matches the document retrieval request from the sorted candidate document list according to the document selection instruction.

[0012] In addition, to achieve the above object, the present application also proposes a personalized document retrieval device, and the personalized document retrieval device includes: A determination module, configured to determine user query information and user position information according to the document retrieval request when receiving a document retrieval request from a user; A generation module, configured to generate personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information; A retrieval module, configured to retrieve relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculate the relevance score between the personalized user query information and each of the relevant documents based on a target scoring model; The determination module is further configured to determine a personalized document that matches the document retrieval request according to the relevance score.

[0013] In addition, to achieve the above object, the present application also proposes a personalized document retrieval device, and the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the personalized document retrieval method as described above.

[0014] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the personalized document retrieval method as described above are implemented.

[0015] One or more technical solutions proposed by this application have at least the following technical effects: when receiving a document retrieval request from a user, determining user query information and user position information according to the document retrieval request; generating personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information; retrieving relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculating a relevance score between the personalized user query information and each of the relevant documents based on a target scoring model; determining a personalized document that matches the document retrieval request according to the relevance score. By the above method, the dynamic changes of user interests are combined with time factors to generate time-sensitive user dynamic interest information, capturing the subtle fluctuations of user interests, and then generating personalized user query information by combining user query information and user position information to more accurately reflect the dynamic needs of users. Then, a lightweight target scoring model is used to calculate the relevance score, which significantly reduces the consumption of computing resources while maintaining the retrieval effect. Then, a personalized document is determined according to the relevance score, thereby effectively improving the efficiency and accuracy of retrieving documents and providing a more fluent interaction experience for users. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of this application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the personalized document retrieval method of this application; Figure 2 It is a schematic flowchart provided for Embodiment 2 of the personalized document retrieval method of this application; Figure 3 It is a schematic module structure diagram of the personalized document retrieval device in the embodiments of this application; Figure 4 It is a schematic device structure diagram of the hardware operating environment involved in the personalized document retrieval method in the embodiments of this application.

[0019] The implementation, functional features, and advantages of the purpose of this application will be further described in combination with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a personalized document retrieval device, etc. The following takes a personalized document retrieval device as an example to illustrate this embodiment and the following embodiments.

[0021] Based on this, the present application embodiment provides a personalized document retrieval method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the personalized document retrieval method of the present application.

[0022] In this embodiment, the personalized document retrieval method includes steps S10 to S40: Step S10: upon receiving a document retrieval request from a user, determining user query information and user position information according to the document retrieval request.

[0023] It should be noted that when a document retrieval request is received from a user, it indicates that personalized document retrieval is required. User query information refers to the initial query information submitted by the user to the personalized document retrieval device. User position information refers to the position information of the user in the enterprise, such as product design engineer, algorithm engineer, etc., which is a relatively stable static information.

[0024] Further, step S10 includes: when receiving a document retrieval request from a user, parsing the document retrieval request; obtaining retrieval request header information according to the parsing result; obtaining retrieval request content according to the parsing result when the verification of the retrieval request header information passes, and determining user query information according to the retrieval request content; obtaining login information of the device that initiated the document retrieval request according to the parsing result; and determining user position information according to the login information.

[0025] It should be understood that the request parsing results of the document retrieval request include but are not limited to the retrieval request header information, the retrieval request content and the login information of the device that initiated the document retrieval request. In order to effectively improve the accuracy of determining the user query information and the user position information, when the retrieval request header information is obtained, it is necessary to verify the retrieval request header information, and determine the user query information and the user position information respectively when the verification passes. Conversely, when the verification fails, the process of determining the user query information and the user position information is directly exited, and a prompt message indicating that the document retrieval request verification failed is displayed.

[0026] It is understandable that after obtaining the content of the retrieval request based on the parsing result, semantic recognition is performed on the content of the retrieval request, and the user's query information is determined according to the semantic recognition result. On the other hand, the login information of the device that initiated the document retrieval request is also extracted from the parsing result. The login information includes, but is not limited to, the login account, login password, and login time, etc. The login account can be the employee number of the user. In an enterprise, the naming of employee numbers for different positions is different. At this time, the user's position information can be determined according to the login information.

[0027] Step S20: Generate personalized user query information according to the user's query information, the user's position information, and the time-sensitive user dynamic interest information.

[0028] It is understandable that the time-sensitive user dynamic interest information is dynamic interest information generated by combining the user's historical behavior information and time factors, and is used to capture the real-time changes of the user's interest. After obtaining the time-sensitive user dynamic interest information, personalized user query information is comprehensively generated in combination with the user's query information and the user's position information. For example, the user's query information is expressed as , the user's position information is expressed as , and the time-sensitive user dynamic interest information is expressed as , then the personalized user query information can be expressed as .

[0029] Step S30: Retrieve relevant documents corresponding to the personalized user query information from the multi-dimensional database, and calculate the correlation score between the personalized user query information and each of the relevant documents based on the target scoring model.

[0030] It should be understood that the multi-dimensional database includes, but is not limited to, vector databases, relational databases, etc. The relevant documents represent a specific text unit in the domain knowledge base corresponding to the personalized user query information, that is, relevant documents corresponding to the personalized user query information are retrieved from the multi-dimensional database according to expert knowledge. The correlation score is used to represent the degree of relevance between the personalized user query information and each relevant document. The larger the value of the correlation score, the more relevant the personalized user query information is to the relevant document.

[0031] Further, before the step of calculating the relevance score between the personalized user query information and each of the relevant documents based on the target scoring model, the following steps are further included: obtaining positive sample documents related to the user query information and negative sample documents not related to the user query information; training the current scoring model based on the pre-trained lightweight re-ranking model, according to the target scoring function, the positive sample documents, and the negative sample documents; calculating the loss value of the current scoring model according to the contrast loss function and the preset temperature parameter; validating and training the current scoring model according to the preset validation samples, and determining the target scoring model when the loss value converges to the preset value.

[0032] It can be understood that in order to effectively improve the accuracy of determining the target scoring model, it is necessary to go through two stages: domain adaptation pre-training and large model-guided fine-tuning training. In the domain adaptation pre-training stage, the adaptability and feature sensitivity of the scoring model to domain data are improved. By constructing positive sample documents related to the user query information and negative sample documents not related to the query information, the scoring model is helped to learn domain knowledge. Among them, positive sample documents refer to documents highly relevant to the query information, and negative sample documents refer to documents not related to the user query information. Through contrast optimization, the pre-trained lightweight re-ranking model can adapt to the characteristics of the domains to which the positive and negative sample documents belong. By using the pre-trained lightweight re-ranking model, the consumption of computing resources can be significantly reduced while maintaining the retrieval effect.

[0033] It should be noted that the specific method for obtaining positive sample documents and negative sample documents can be as follows: The preferred document storage method is to use a combination of a vector database and Elasticsearch (ES) for storage. When retrieving, the text of the user query can be vectorized and matched with the document vectors in the vector database. At the same time, the full-text retrieval ability of ES is used to retrieve the document content. Finally, by combining the retrieval results of both, a more accurate and comprehensive retrieval result is obtained, and the documents in the retrieval result are used as positive sample documents, while negative sample documents are randomly selected from the set of documents that have never been retrieved, that is, documents that do not appear in the retrieval result are used as negative sample documents, and the pre-trained model is used to distinguish the differences between the sample documents and the negative sample documents.

[0034] It should be understood that in order to improve the accuracy of the pre-trained model, a pre-trained lightweight re-ranking model is used as the base network for training the current scoring model, so as to utilize the feature representation ability accumulated by the pre-trained model on the general dataset, thereby enabling a basic ranking ability for the retrieval task. The target scoring function refers to the function used to measure the relevance between a document and query information. The higher the measurement score, the more relevant it indicates. Therefore, the relevance between the positive sample document and the user query information is much greater than the relevance between the negative sample document and the user query information. After training the current scoring model, the loss value of the current scoring model is calculated according to the contrast loss function and the preset temperature parameter. The contrast loss function can be expressed as: 。

[0035] Among them, represents the contrast loss function, represents maximizing the relevance between the positive sample document and the user query information, represents minimizing the relevance between the negative sample document and the user query information, represents the preset temperature parameter.

[0036] Furthermore, the steps of determining the target scoring model when the loss value converges to a preset value include: when the loss value converges to the preset value, obtaining the document retrieval results output by the target large model; performing semantic analysis on each document in the document retrieval results, and performing relevance analysis on each document after semantic analysis to generate a personalized document ranking list; counting the number of documents in the personalized document ranking list, and constructing a target document pair according to the number; fine-tuning and training the verified and trained current scoring model according to the target document pair to obtain the target scoring model.

[0037] It should be noted that the target scoring model has the ability of multilingual retrieval and can effectively score the relevance between personalized user query information and each relevant document. In order to enhance the understanding and processing ability of the scoring model for domain-specific data, when the loss value converges to the preset value, a large model-based evaluation and correction framework can also be used to fine-tune and train the verified and trained current scoring model. Specifically: performing semantic analysis on each document in the document retrieval results, and performing relevance analysis on each document after semantic analysis. For two documents, the score of the scoring model for the document in the retrieval results is expressed as That is, , the greater the difference in scores is from the true sequence, the greater the loss. Compared with the traditional binary logic loss function, the ranking loss function can effectively capture the relative relevance between queries and documents and optimize the relative ranking of document pairs. Because this ranking loss function not only focuses on the absolute scores of documents but also emphasizes the relative ranking of different document pairs, enabling the model to better learn the subtle differences in distinguishing high and low relevance in complex ranking tasks. This ranking loss function can be the RankNet function.

[0038] It should be understood that after generating the personalized document sorted list, target document pairs are constructed according to the number of documents in the personalized document sorted list. For example, if M represents the number of documents in the personalized document sorted list, the number of target document pairs can be M(M - 1) / 2, which can provide rich documents for fine-tuning training. Then, the current scoring model after validation training is fine-tuned according to the target document pairs. Through fine-tuning training, the current scoring model after validation training can more accurately calculate the relevance scores between personalized user query information and each relevant document, and make full use of the knowledge of the large model to effectively improve the robustness of the target scoring model when dealing with tasks in the field of document retrieval.

[0039] Step S40: Determine personalized documents that match the document retrieval request according to the relevance scores.

[0040] It can be understood that personalized documents refer to the documents that best match the document retrieval request and meet the personalized needs of users. After obtaining the relevance scores, personalized documents that match the document retrieval request are selected from the sorted candidate document list according to the relevance scores.

[0041] Furthermore, step S40 includes: generating a candidate document list according to each relevant document corresponding to the relevance score; performing intelligent sorting on the candidate document list according to the relevance score; and in response to a document selection instruction triggered by the user on the document display page, and selecting a personalized document that matches the document retrieval request from the sorted candidate document list according to the document selection instruction.

[0042] It should be understood that after generating the candidate document list according to each relevant document corresponding to the relevance score, the relevance scores are intelligently sorted according to the relevance score to ensure that the most relevant candidate documents are ranked at the front. In this way, users can preferentially view the documents that are most relevant to the query information. Since the final choice is still in the hands of the user, at this time, it is necessary to respond to the document selection instruction triggered by the user on the document display page and select the document corresponding to the document selection instruction from the sorted candidate document list, which is the personalized document that matches the document retrieval request, thereby effectively improving the user experience.

[0043] In this embodiment, when a document retrieval request from a user is received, user query information and user position information are determined according to the document retrieval request; personalized user query information is generated according to the user query information, the user position information, and time-sensitive user dynamic interest information; relevant documents corresponding to the personalized user query information are retrieved from a multi-dimensional database, and a relevance score between the personalized user query information and each of the relevant documents is calculated based on a target scoring model; and a personalized document that matches the document retrieval request is determined according to the relevance score. By the above method, the dynamic change of user interest is combined with time factors to generate time-sensitive user dynamic interest information, capturing the subtle fluctuations of user interest, and then personalized user query information is generated by combining the user query information and the user position information to more accurately reflect the dynamic needs of the user. Then, a lightweight target scoring model is used to calculate the relevance score, so that while maintaining the retrieval effect, the consumption of computing resources is significantly reduced. Then, a personalized document is determined according to the relevance score, thereby effectively improving the efficiency and accuracy of retrieving documents and providing a more fluent interaction experience for the user.

[0044] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S20 includes steps S201 to S204: Step S201, obtaining the historical behavior information of the user from the target historical record storage database.

[0045] It should be noted that the historical behavior information of the user is stored in the target historical record storage database and can be stored in the form of a hash table, including a JSON-formatted structured object of the user ID and the user behavior record, and each object contains fields such as a timestamp, a query content, and click document information. In the target historical record storage database, an expiration policy is adopted to clean up expired data regularly to ensure that the historical behavior information obtained from the target historical record storage database is recent.

[0046] Step S202, extracting keywords from the historical behavior information based on a professional concept dictionary.

[0047] It can be understood that in order to effectively improve the accuracy of keyword extraction, a professional concept dictionary is introduced. Since the dynamic interest representation depending on the task focuses on the real-time change of user behavior, therefore, historical behavior keywords are extracted from the historical behavior information based on the professional concept dictionary, which can deeply explore the interest preferences of the user in different time periods.

[0048] Step S203: Generate time-sensitive user dynamic interest information based on historical behavior keywords.

[0049] Further, step S203 includes: setting an initial weight for the historical behavior keywords; obtaining the number of occurrences of the historical behavior keywords; adjusting the initial weight according to the number of occurrences and a target decay factor to obtain a target weight; sorting the historical behavior keywords according to the target weight, and generating time-sensitive user dynamic interest information according to the sorting result.

[0050] It can be understood that since the importance of historical behavior keywords in the user's historical records will change continuously, a decay mechanism is introduced. As time goes by or the user's interests shift, the weights of historical behavior keywords will gradually decrease, so as to ensure that the personalized document retrieval device gives priority to recently important and highly relevant historical behavior keywords when making recommendations. For this reason, an initial weight is first set for the historical behavior keywords, and this initial weight can be 1. For example, when the keyword appears for the first time, , for subsequent occurrences of the vocabulary, the weight of the historical behavior keyword will be calculated by decay according to the number of occurrences, that is .

[0051] Among them, represents the target decay factor.

[0052] It should be understood that after obtaining the target weight using the above formula, sorting the historical behavior keywords according to the target weight, and generating time-sensitive user dynamic interest information according to the sorting result can be expressed as , where represents the historical behavior keyword, represents the target weight, which characterizes the relative importance of this historical behavior keyword in the user's historical behavior.

[0053] Step S204: Generate personalized user query information according to the user query information, the user position information, and the time-sensitive user dynamic interest information.

[0054] It can be understood that after obtaining the time-sensitive user dynamic interest information, personalized user query information is generated by combining the user query information and the user position information. For example, the user query information is expressed as , the user position information is expressed as , and the time-sensitive user dynamic interest information is expressed as , then the personalized user query information can be expressed as .

[0055] In this embodiment, historical behavior information of a user is obtained from a target historical record storage database; keyword extraction is performed on the historical behavior information based on a professional concept dictionary; time-sensitive user dynamic interest information is generated according to the historical behavior keywords; and personalized user query information is generated according to the user query information, the user position information, and the time-sensitive user dynamic interest information. By the above method, after obtaining the historical behavior information of the user, historical behavior keywords are extracted from the historical behavior information based on the professional concept dictionary, so as to generate time-sensitive user dynamic interest information by combining the historical behavior keywords and time factors, achieving the purpose of characterizing the time adaptability of the user's interests, and then personalized user query information is generated by combining the user query information and the user position information, thereby effectively improving the accuracy of generating personalized user query information.

[0056] This application also provides a personalized document retrieval device. Please refer to Figure 3 , and the personalized document retrieval device includes: A determination module 10, configured to determine user query information and user position information according to the document retrieval request when receiving the document retrieval request of the user.

[0057] A generation module 20, configured to generate personalized user query information according to the user query information, the user position information, and the time-sensitive user dynamic interest information.

[0058] A retrieval module 30, configured to retrieve relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculate a relevance score between the personalized user query information and each of the relevant documents based on a target scoring model.

[0059] The determination module 10 is further configured to determine a personalized document that matches the document retrieval request according to the relevance score.

[0060] This embodiment determines user query information and user position information according to the document retrieval request when receiving a document retrieval request from the user; generates personalized user query information according to the user query information, the user position information and time-sensitive user dynamic interest information; retrieves relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculates the correlation score between the personalized user query information and each of the relevant documents based on a target scoring model; and determines personalized documents that match the document retrieval request according to the correlation score. In the above manner, the dynamic changes of user interests are combined with time factors to generate time-sensitive user dynamic interest information, capture subtle fluctuations in user interests, and then combine user query information and user position information to generate personalized user query information so as to more accurately reflect the user's dynamic needs, and then use a lightweight target scoring model to calculate the correlation score, so as to significantly reduce the consumption of computing resources while maintaining the retrieval effect, and then determine personalized documents according to the correlation score, thereby effectively improving the efficiency and accuracy of document retrieval and providing users with a smoother interactive experience.

[0061] The personalized document retrieval device provided by the present application adopts the personalized document retrieval method in the above embodiment, which can solve the technical problem of low efficiency and accuracy of document retrieval in the prior art. Compared with the prior art, the beneficial effects of the personalized document retrieval device provided by the present application are the same as the beneficial effects of the personalized document retrieval method provided by the above embodiment, and other technical features in the personalized document retrieval device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0062] In one embodiment, the determination module 10 is also used to parse the document retrieval request of the user when receiving the document retrieval request; obtain the retrieval request header information according to the parsing result; obtain the retrieval request content according to the parsing result when the verification of the retrieval request header information passes, and determine the user query information according to the retrieval request content; obtain the login information of the device that initiated the document retrieval request according to the parsing result; and determine the user position information according to the login information.

[0063] In one embodiment, the generation module 20 is also used to obtain the user's historical behavior information from a target history record storage database; perform keyword extraction on the historical behavior information based on a professional concept dictionary; generate time-sensitive user dynamic interest information based on historical behavior keywords; and generate personalized user query information based on the user query information, the user position information, and the time-sensitive user dynamic interest information.

[0064] In one embodiment, the generation module 20 is further configured to set an initial weight for the historical behavior keywords; obtain the number of occurrences of the historical behavior keywords; adjust the initial weight according to the number and a target decay factor to obtain a target weight; sort the historical behavior keywords according to the target weight, and generate time-sensitive user dynamic interest information according to the sorting result.

[0065] In one embodiment, the retrieval module 30 is further configured to obtain positive sample documents related to the user query information and negative sample documents unrelated to the user query information; train a current scoring model based on a pre-trained lightweight re-ranking model, according to a target scoring function, the positive sample documents, and the negative sample documents; calculate a loss value of the current scoring model according to a contrastive loss function and a preset temperature parameter; perform validation training on the current scoring model according to preset validation samples, and determine a target scoring model when the loss value converges to a preset value.

[0066] In one embodiment, the retrieval module 30 is further configured to, when the loss value converges to a preset value, obtain a document retrieval result output by a target large model; perform semantic analysis on each document in the document retrieval result, and perform relevance analysis on each semantically analyzed document to generate a personalized document sorted list; count the number of documents in the personalized document sorted list, and construct a target document pair according to the number; perform fine-tuning training on the current scoring model after validation training according to the target document pair to obtain a target scoring model.

[0067] In one embodiment, the determination module 10 is further configured to generate a candidate document list according to each relevant document corresponding to the relevance score; perform intelligent sorting on the candidate document list according to the relevance score; respond to a document selection instruction triggered by a user on a document display page, and select a personalized document that matches the document retrieval request from the sorted candidate document list according to the document selection instruction.

[0068] The present application provides a personalized document retrieval device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the personalized document retrieval method in the first embodiment above.

[0069] Next, refer to Figure 4, which shows a schematic structural diagram of a personalized document retrieval device suitable for implementing the embodiments of the present application. The personalized document retrieval device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player: portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown personalized document retrieval device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0070] As Figure 4 shown, the personalized document retrieval device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a ROM (Read Only Memory) 1002 or a program loaded from a storage device 1003 into a RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the personalized document retrieval device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the personalized document retrieval device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a personalized document retrieval device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0071] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. The computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the disclosed embodiments of the present application are executed.

[0072] The personalized document retrieval device provided in the present application adopts the personalized document retrieval method in the above embodiment, and can solve the technical problem of low efficiency and accuracy in retrieving documents in the prior art. Compared with the prior art, the beneficial effects of the personalized document retrieval device provided in the present application are the same as those of the personalized document retrieval method provided in the above embodiment, and other technical features in the personalized document retrieval device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0073] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0074] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0075] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the personalized document retrieval method in the above embodiment.

[0076] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0077] The above computer-readable storage medium can be included in a personalized document retrieval device; or it can exist separately without being assembled into the personalized document retrieval device.

[0078] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0080] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0081] The readable storage medium provided by the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned personalized document retrieval method, and can solve the technical problems of low efficiency and accuracy in retrieving documents in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the personalized document retrieval method provided by the above embodiments, and will not be elaborated here.

[0082] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A personalized document retrieval method, characterized in that, The method includes: When receiving a document retrieval request from a user, determining user query information and user position information according to the document retrieval request; Generating personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information; Retrieving relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculating a relevance score between the personalized user query information and each of the relevant documents based on a target scoring model; Determining personalized documents that match the document retrieval request according to the relevance score.

2. The method according to claim 1, wherein The step of generating personalized user query information according to the user query information, the user position information, and time-sensitive user dynamic interest information includes: Obtaining the historical behavior information of the user from a target historical record storage database; Performing keyword extraction on the historical behavior information based on a professional concept dictionary; Generating time-sensitive user dynamic interest information according to historical behavior keywords; Generating personalized user query information according to the user query information, the user position information, and the time-sensitive user dynamic interest information.

3. The method according to claim 2, wherein The step of generating time-sensitive user dynamic interest information according to historical behavior keywords includes: Setting an initial weight for the historical behavior keywords; Obtaining the number of occurrences of the historical behavior keywords; Adjusting the initial weight according to the number and a target decay factor to obtain a target weight; Sorting the historical behavior keywords according to the target weight, and generating time-sensitive user dynamic interest information according to the sorting result.

4. The method according to claim 1, characterized in that, Before the step of calculating a relevance score between the personalized user query information and each of the relevant documents based on a target scoring model, it further includes: Obtaining positive sample documents related to the user query information and negative sample documents unrelated to the user query information; Training a current scoring model based on a pre-trained lightweight re-ranking model, according to a target scoring function, the positive sample documents, and the negative sample documents; Calculating a loss value of the current scoring model according to a contrastive loss function and a preset temperature parameter; Performing validation training on the current scoring model according to a preset validation sample, and determining a target scoring model when the loss value converges to a preset value.

5. The method according to claim 4, wherein The step of determining a target scoring model when the loss value converges to a preset value includes: When the loss value converges to a preset value, obtaining a document retrieval result output by a target large model; Performing semantic analysis on each document in the document retrieval result, and performing relevance analysis on each document after semantic analysis to generate a personalized document sorted list; Counting the number of documents in the personalized document sorted list, and constructing a target document pair according to the number; Performing fine-tuning training on the current scoring model after validation training according to the target document pair to obtain a target scoring model.

6. The method according to claim 1, characterized in that, The step of determining user query information and user position information according to the document retrieval request when receiving a document retrieval request from a user includes: When receiving a document retrieval request from a user, parsing the document retrieval request; Get the retrieval request header information according to the parsing result; When the verification of the search request header information is passed, the search request content is obtained according to the parsing result, and the user query information is determined according to the search request content; Obtaining login information of a device that initiates the document retrieval request according to the parsing result; The user's position information is determined according to the login information.

7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the personalized document that matches the document retrieval request according to the relevance score includes: generating a candidate document list according to each relevant document corresponding to the relevance score; Intelligently sorting the candidate document list according to the relevance score; In response to a document selection instruction triggered by a user on a document display page, a personalized document matching the document retrieval request is selected from a sorted candidate document list according to the document selection instruction.

8. A personalized document retrieval device, characterized in that, The device comprises: A determination module, configured to determine user query information and user position information according to the document retrieval request when receiving a document retrieval request from the user; A generating module, configured to generate personalized user query information according to the user query information, the user position information and time-sensitive user dynamic interest information; A retrieval module, configured to retrieve relevant documents corresponding to the personalized user query information from a multi-dimensional database, and calculate a correlation score between the personalized user query information and each of the relevant documents based on a target scoring model; The determination module is further configured to determine a personalized document that matches the document retrieval request according to the relevance score.

9. A personalized document retrieval device, characterized in that, The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the personalized document retrieval method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the personalized document retrieval method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Document sorting method and device, electronic equipment and storage medium

    CN113032549A

  • Medical question answering processing method and device based on network resources

    CN114780672A

  • Retrieval enhancement method and device based on large model and storage medium

    CN119988602A

  • Intelligent semantic document recommendation method and device, and computer-readable storage medium

    WO2020224097A1

Cited By

  • Knowledge base document searching method and device, electronic equipment and storage medium

    CN120821830A