Semantic search method, system and equipment based on fine tuning pre-training model and medium
By adopting a semantic search method based on fine-tuning pre-trained models in information retrieval, the problem of insufficient text semantic understanding ability in the prior art is solved, and more accurate and efficient search results are achieved, improving user experience and model adaptability.
Patent Information
- Application Number
- CN202510337902.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively understand the deep semantics of text in information retrieval, resulting in inaccurate or uncorrelated search results. The method based on simple neural network has limited semantic understanding capabilities, requires a large amount of labeled data, and has poor adaptability.
The semantic search method based on fine-tuning pre-trained model is adopted, and the basic data set is fine-tuned by selecting the chinese-macbert-base model in the Sentence Transformer architecture, and the model is optimized using infoNCE loss and PEFT strategy, combined with vector database efficient retrieval technology to process user query requests, and dynamically update the model according to user behavior.
It significantly improves the model's understanding of text semantics and the accuracy of similarity calculation, improves the click-through rate of user search query, enhances the relevance and accuracy of search results, reduces dependence on a large amount of labeled data, and improves the model's adaptability and search speed.
Smart Images

Figure CN120218081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a semantic search method, system, device and medium based on a fine-tuned pre-trained model. Background Art
[0002] With the rapid development of information technology and the popularization of the Internet, the amount of information has shown an explosive growth. People urgently need to quickly obtain valuable information from massive data through search engines or recommendation systems. Traditional keyword-based search methods are effective in dealing with small-scale information, but with the explosion of information volume and content diversification, their limitations have gradually emerged, resulting in inaccurate or irrelevant search results. Therefore, improving search quality, reducing costs, and enhancing user experience have become important topics in the field of information retrieval.
[0003] To solve this problem, semantic search has emerged. It uses deep learning and natural language processing technology (NLP) to understand and reason about the deep meaning of text, thereby improving the accuracy and quality of search. Among them, semantic similarity search, as an important part, plays a key role in application scenarios such as information retrieval, recommendation systems, and question answering systems. Through deep semantic analysis, it can effectively identify and match the content most relevant to user needs, improving the quality of search results.
[0004] In search technology, existing methods include rule-based information retrieval, text-based information retrieval, and machine learning-based information retrieval, among which: Rule-based search methods rely on manually set rules for keyword or pattern matching, lacking flexibility and adaptability, and it is difficult to handle new language expressions or unforeseen search needs. Text-based search methods, especially inverted index and keyword matching, although having high retrieval efficiency, cannot understand the deep semantic relationships between words, resulting in the "keyword island" problem, affecting the accuracy and quality of search results. Search methods based on simple neural networks and machine learning solve the limitations of keyword matching to a certain extent, but their semantic understanding ability is limited. They usually can only capture the surface patterns of text and require a large amount of labeled data, and have poor adaptability to small data sets.
[0005] Therefore, how to improve the accuracy of semantic matching while ensuring high efficiency has become the core issue in the current research of information retrieval technology. Summary of the Invention
[0006] The object of the present invention is to provide a semantic search method, system, device and medium based on a fine-tuned pre-trained model, so as to solve all or one of the above problems existing in the prior art.
[0007] To solve the above technical problems, the specific technical solution of the present invention is as follows: On the one hand, the present invention provides a semantic search method based on a fine-tuned pre-trained model, including the following steps: Model offline training step: Create a basic dataset, and perform parameter fine-tuning training on the basic model based on the basic dataset; Online inference step: In response to a new query request from a user, process the new query request based on the model after parameter fine-tuning training by using a vector database efficient retrieval technology, and output a requirement document related to the new query request; Model update step: Dynamically update the model according to the dynamic changes of user behavior data.
[0008] As an improved solution, before performing parameter fine-tuning training on the basic model based on the basic dataset, it includes: Select the chinese-macbert-base model in the Sentence Transformer architecture as the basic model.
[0009] As an improved solution, the parameter fine-tuning training on the basic model based on the basic dataset includes: Perform training on the basic model based on the basic dataset; During the training process, perform model parameter fine-tuning based on the infoNCE loss; The strategy of parameter fine-tuning includes: the PEFT strategy.
[0010] As an improved solution, the steps of the parameter fine-tuning training include: Call the basic model to extract query and doc_title from the basic dataset; Call the basic model to calculate the embedding representations of query and doc_title, and obtain query_embedding and doc_title_embedding; Call the basic model to calculate the similarity between query_embedding and doc_title_embedding, and obtain similaty; Call the basic model to calculate loss according to similaty and label; Update the model parameters of the basic model based on the optimizer.
[0011] As an improved solution, the model after fine-tuning training based on the parameters processes the new query request by using the vector database efficient retrieval technology, and outputs the requirement document related to the new query request, including: Call the model after fine-tuning training to convert the new query request into a vector form, and output the user query vector; Call the model after fine-tuning training to calculate the similarity between the user query vector and the document vectors stored in the vector database; Output the document vector in the vector database with the highest similarity to the user query vector as the result to the user.
[0012] As an improved solution, before the model after fine-tuning training based on the parameters processes the new query request by using the vector database efficient retrieval technology and outputs the requirement document related to the new query request, it further includes: In response to the training stage of the base model, the base model maps the user's query and documents into a common feature space, and evaluates the semantic similarity between the user's query and documents; In response to the end of the training of the base model, the base model performs vectorization processing on all documents, and stores the results of the vectorization processing in the vector database.
[0013] As an improved solution, the dynamic update of the model according to the dynamic changes of user behavior data includes: Periodically update the model, and compare the updated model with the old model on the test set; In response to the updated model performing better than the old model, deploy the updated model to the online environment; In response to the old model performing better than the updated model, keep using the old model.
[0014] On the other hand, the present invention also provides a semantic search system based on a fine-tuned pre-trained model, including: A model offline training module, which is used to: create a basic data set, and perform parameter fine-tuning training on a base model based on the basic data set; An online inference module, which is used to: in response to a new query request from a user, process the new query request by using the vector database efficient retrieval technology based on the model after fine-tuning training, and output a requirement document related to the new query request; A model update module, which is used to: dynamically update the model according to the dynamic changes of user behavior data.
[0015] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the semantic search method based on the fine-tuned pre-trained model are implemented.
[0016] On the other hand, the present invention also provides a computer device, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; wherein: The memory is used to store a computer program; The processor is used to execute the steps of the semantic search method based on the fine-tuned pre-trained model by running the program stored on the memory.
[0017] The beneficial effects of the technical solution of the present invention are: 1. The semantic search method based on the fine-tuned pre-trained model of the present invention can significantly improve the model's ability to understand text semantics and the accuracy of similarity calculation through fine-tuning on domain-specific labeled data. It not only improves the click-through rate (CTR) of the user's search query, making the search results more relevant and accurate, enhancing user satisfaction and search efficiency, but also optimizes the training process from the perspective of model training. By using the pre-trained model, it reduces the consumption of training time and computing resources, reduces the dependence on a large amount of labeled data, and enables the model to quickly adapt to specific search tasks; by introducing the FAISS database, it further significantly improves the search speed in search inference, can process a large number of user requests in a short time, reduces the user waiting time, and improves the response speed and user experience of the search service.
[0018] 2. The semantic search system based on the fine-tuned pre-trained model of the present invention can, through the mutual cooperation of system modules, implement the semantic search method based on the fine-tuned pre-trained model of the present invention.
[0019] 3. The computer-readable storage medium of the present invention can guide the system modules to cooperate, thereby implementing the semantic search method based on the fine-tuned pre-trained model of the present invention, and the computer-readable storage medium of the present invention also effectively improves the operability of the semantic search method based on the fine-tuned pre-trained model.
[0020] 4. The computer device of the present invention can store and execute the computer-readable storage medium, thereby implementing the semantic search method based on the fine-tuned pre-trained model of the present invention. Description of the Drawings
[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic flowchart of the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 of the present invention; Figure 2 It is a schematic diagram of the training architecture of the model in the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 of the present invention; Figure 3 It is a schematic diagram of the inference architecture of the model in the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 of the present invention; Figure 4 It is a schematic diagram of the update architecture of the model in the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 of the present invention; Figure 5 It is a schematic diagram of the architecture of the semantic search system based on the fine-tuned pre-trained model described in Embodiment 2 of the present invention; Figure 6 It is a schematic diagram of the structure of the computer device described in Embodiment 4 of the present invention; The label descriptions in the accompanying drawings are as follows: 1501, processor; 1502, communication interface; 1503, memory; 1504, communication bus. Specific embodiments
[0023] The following will elaborate on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0024] In the description of the present invention, it should be noted that the embodiments described in the present invention are some embodiments of the present invention, rather than all embodiments; all other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the protection scope of the present invention.
[0025] In the description, claims and the above drawings of this document, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this document described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment. Embodiment 1
[0026] This embodiment provides a semantic search method based on fine-tuning a pre-trained model, as Figures 1 to 4 shown, including the following steps: S100. Model offline training step, including: S101. Dataset acquisition: Obtain the historical behavior data of users on the Internet platform as the basic dataset.
[0027] S102. Training and fine-tuning: Select the chinese-macbert-base model in the Sentence Transformer architecture as the basis (this model is based on BERT technology and is suitable for processing sentence-level semantic tasks); perform model training based on the basic dataset and perform model fine-tuning based on the infoNCE loss; for the fine-tuning technology, parameter-efficient fine-tuning (PEFT) can be selected to improve the performance of the pre-trained model on new tasks by minimizing the number of fine-tuned parameters and the computational complexity, and finally enable the model to distinguish positive and negative samples and enhance the model's semantic similarity discrimination ability; Specifically, the training architecture is as Figure 2 shown, and the training steps are as follows: (1) The model extracts query and doc_title from the basic dataset, calculates their embedding representations, and obtains query_embedding and doc_title_embedding; (2) The model calculates the similarity between query_embedding and doc_title_embedding to obtain similaty; (3) The model calculates the loss based on similaty and label, and updates the model parameters based on the optimizer.
[0028] S103. Data Vectorization: Additionally, during the model training phase, the model learns to map the user's queries and documents into a common feature space to facilitate the evaluation of the semantic similarity between them. After training, the model vectorizes all the documents, converts them into vector representations, and stores them.
[0029] S200. Online Inference Steps, including: S201. Vector Calculation: In response to a new query request initiated by the user, the request is input into the model, which converts it into vector form and outputs a vector representation corresponding to the query semantics, i.e., the user query vector.
[0030] S202. Document Retrieval: The model calculates the similarity between the user query vector and the document vectors stored in the FAISS database; FAISS quickly returns the document vectors similar to the query vector, sorts these document vectors according to the similarity scores of the document vectors returned by FAISS, and outputs the most relevant (highest similarity) documents to the user. Its specific inference architecture is as Figure 3 shown; It should be noted that the FAISS database can quickly retrieve the vector most similar to a given vector in a large-scale dataset; to achieve this effect, an index needs to be built for all document vectors in the FAISS database in advance; building the index is a one-time process, but it needs to be updated regularly to include new document vectors; the FAISS database has higher query efficiency than traditional databases, and it is specifically optimized for vector similarity search, which greatly improves the efficiency of this method and enhances the user experience and retrieval efficiency.
[0031] S300. Model Update Steps, including: S301. To adapt to the dynamic changes of user behavior data, the model is updated regularly to maintain the recommendation effect; Specifically, the model is updated every day and compared with the old model on the test set; when the new model performs better, the new model is deployed to the online environment; when the old model performs better, the old model continues to be used to ensure the stability of the online service and the recommendation quality. The specific model update architecture is as Figure 4 shown.
[0032] It should be noted that the above examples are only for explaining the present invention and should not be used to limit the protection scope of the present invention. Embodiment 2
[0033] This embodiment is based on the same inventive concept as the semantic search method based on fine-tuning a pre-trained model described in Embodiment 1, and provides a semantic search system based on fine-tuning a pre-trained model, as Figure 5 shown, including: The model offline training module is used to: create a basic dataset and perform parameter fine-tuning training on the basic model based on the basic dataset; The online inference module is used to: in response to a new query request from a user, process the new query request based on the model after parameter fine-tuning training by using the vector database efficient retrieval technology, and output a requirement document related to the new query request; The model update module is used to: dynamically update the model according to the dynamic changes of user behavior data. Embodiment 3
[0034] This embodiment provides a computer-readable storage medium, including: The storage medium is used to store computer software instructions for implementing the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 above, and it includes a program set for executing the semantic search method based on the fine-tuned pre-trained model; specifically, the executable program can be built into the semantic search system based on the fine-tuned pre-trained model described in Embodiment 2. In this way, the semantic search system based on the fine-tuned pre-trained model can implement the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 by executing the built-in executable program.
[0035] In addition, the computer-readable storage medium of this embodiment can adopt any combination of one or more readable storage media, where the readable storage medium includes an electrical, optical, electromagnetic, infrared or semiconductor system, device or component, or any combination of the above. Embodiment 4
[0036] This embodiment provides an electronic device, as Figure 6 shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504. Among them, the processor 1501, the communication interface 1502, and the memory 1503 complete mutual communication through the communication bus 1504.
[0037] The memory 1503 is used to store a computer program; The processor 1501 is used to implement the steps of the semantic search method based on the fine-tuned pre-trained model described in Embodiment 1 above when executing the computer program stored on the memory 1503.
[0038] As an implementation manner of the present invention, the communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 6 only a thick line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus.
[0039] As an implementation manner of the present invention, the communication interface is used for communication between the above terminal and other devices.
[0040] As an implementation manner of the present invention, the memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0041] As an implementation manner of the present invention, the above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0042] Different from the prior art, by using a semantic search method, system, device and medium based on fine-tuning a pre-trained model according to this application, the understanding ability of the model for text semantics and the accuracy of similarity calculation can be significantly improved through fine-tuning on domain-specific labeled data. This not only improves the click-through rate (CTR) of the user search query, making the search results more relevant and accurate, enhancing user satisfaction and search efficiency, but also optimizes the training process from the perspective of model training. The pre-trained model is used to reduce the consumption of training time and computing resources, reduce the dependence on a large amount of labeled data, and enable the model to quickly adapt to specific search tasks. By introducing the FAISS database, a significant improvement in search speed is achieved in search inference, enabling a large number of user requests to be processed in a short time, reducing the user waiting time, and improving the response speed of the search service and the user experience.
[0043] It should be understood that in various embodiments of this article, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this article.
[0044] It should also be understood that in the embodiments of this article, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects.
[0045] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.
[0046] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0047] In several embodiments provided in this document, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.
[0048] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the objectives of the embodiments of this document.
[0049] Furthermore, in each embodiment of this document, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0050] If the above-mentioned integrated units are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this document. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0051] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. All equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.
Claims
1. A semantic search method based on fine-tuning a pre-trained model, characterized in that: The following steps are involved: Model offline training steps: Creating a basic data set, and performing parameter fine-tuning training on a basic model based on the basic data set; Online reasoning steps: In response to a new query request from a user, the model trained by fine-tuning the parameters is used to process the new query request using a vector database efficient retrieval technology, and a requirement document related to the new query request is output; Model update steps: The model is dynamically updated according to the dynamic changes of user behavior data.
2. The semantic search method based on fine-tuning the pre-training model according to claim 1, characterized in that: Before fine-tuning the parameters of the basic model based on the basic data set, the method includes: The chinese-macbert-base model in the Sentence Transformer architecture is selected as the base model.
3. The semantic search method based on fine-tuning the pre-training model according to claim 1, characterized in that: The fine-tuning of parameters of the basic model based on the basic data set includes: Training the basic model based on the basic data set; During the training process, model parameters are fine-tuned based on infoNCE loss; Strategies for parameter fine-tuning include: PEFT strategy.
4. The semantic search method based on fine-tuning the pre-training model according to claim 2, characterized in that: The step of parameter fine-tuning training includes: Call the basic model to extract query and doc_title from the basic data set; Call the basic model to calculate the embedding representation of query and doc_title to obtain query_embedding and doc_title_embedding; Call the basic model to calculate the similarity between query_embedding and doc_title_embedding to obtain similarity; Call the basic model to calculate the loss according to similarity and label; The model parameters of the base model are updated based on the optimizer.
5. The semantic search method based on fine-tuning the pre-training model according to claim 1, characterized in that: The model trained based on the parameter fine-tuning uses a vector database efficient retrieval technology to process the new query request, and outputs a requirement document related to the new query request, including: Calling the model trained by parameter fine-tuning to convert the new query request into a vector form, and outputting a user query vector; Calling the model trained by parameter fine-tuning to calculate the similarity between the user query vector and the document vector stored in the vector database; The document vector in the vector database having the highest similarity to the user query vector is output to the user as a result.
6. The semantic search method based on fine-tuning the pre-training model according to claim 5, characterized in that: Before processing the new query request based on the model trained by parameter fine-tuning using the vector database efficient retrieval technology and outputting the requirement document related to the new query request, the method further includes: In response to the training phase of the base model, the base model maps the user's query and the document into a common feature space and evaluates the semantic similarity between the user's query and the document; In response to the completion of the training of the basic model, the basic model vectorizes all documents and stores the vectorized results in a vector database.
7. The semantic search method based on fine-tuning the pre-training model according to claim 1, characterized in that: The dynamically updating the model according to the dynamic changes of the user behavior data includes: Update the model periodically and compare the updated model with the old model on the test set; In response to the updated model performing better than the old model, deploying the updated model in an online environment; In response to the old model outperforming the updated model, use of the old model is maintained.
8. A semantic search system based on a fine-tuned pre-trained model, characterized in that: include: The model offline training module is used to: create a basic data set, and perform parameter fine-tuning training on the basic model based on the basic data set; The online reasoning module is used to: respond to a new query request from a user, process the new query request using a vector database efficient retrieval technology based on the model trained by parameter fine-tuning, and output a requirement document related to the new query request; The model updating module is used to dynamically update the model according to the dynamic changes of user behavior data.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the semantic search method based on fine-tuning the pre-trained model described in any one of claims 1 to 7 are implemented.
10. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein: The memory is used to store computer programs; The processor is used to execute the steps of the semantic search method based on fine-tuning the pre-trained model described in any one of claims 1 to 7 by running the program stored in the memory.