Searching method and device, equipment and medium

By performing large-scale text conversion of query text and converting non-standard vocabulary into standard vocabulary, the low correlation and accuracy problems caused by the language differences between the problem text and database text are solved, and higher quality search results are achieved.

CN120067280APending Publication Date: 2025-05-30BEIJING LUYE NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150044.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, there is a language difference between the problem text being too colloquial and the text stored in the database, resulting in low correlation between the searched related text and the problem text and low search accuracy.

Method used

By calling the fine-tuned big model to convert the query text, convert non-standard vocabulary into standard vocabulary, obtain the converted target query text, and find the target text that matches the target query text in the database.

Benefits of technology

Enhanced the matching degree between the target query text and the text stored in the database, improve the quality of search results, and enables the desired target text to be found faster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067280A_ABST
    Figure CN120067280A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a search method and device, equipment and a medium. In the embodiment of the invention, an input search instruction is received, and a query text carried in the search instruction is obtained; calling the fine-tuned large model to perform text conversion on the query text, and converting non-standard vocabularies in the query text into standard vocabularies to obtain a converted target query text; and searching a target text matched with the target query text in the database and outputting the target text. In the embodiment of the invention, the query text is polished through the fine-tuned large model to obtain the target query text carrying the standard vocabularies, so that the matching degree between the target query text and the text stored in the database is enhanced, the required target text can be found more quickly, the quality of the search result is obviously improved, and the search efficiency is improved. By enhancing the matching degree between the user query and the document content, the search experience is effectively improved, and more accurate and personalized information service is provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a search method, device, equipment, and medium. Background Art

[0002] With the development of Internet technology, search engines have become an important tool for people to obtain information. Users input question texts, and the search engines search for relevant texts matching the question texts in the database and output them.

[0003] In related technologies, search engines mainly rely on word matching based on word segmentation and semantic matching technology based on text embedding to search for relevant texts matching the question texts in the database. However, the question texts may be too colloquial, and there are language differences between the question texts and the texts stored in the database, resulting in low relevance between the finally searched relevant texts and the question texts and low search accuracy. Summary of the Invention

[0004] This application provides a search method, device, equipment, and medium to solve the problem that there are language differences between the existing overly colloquial question texts and the texts stored in the database, resulting in low relevance between the finally searched relevant texts and the question texts and low search accuracy.

[0005] In a first aspect, an embodiment of this application provides a search method, and the method includes:

[0006] Receive an input search instruction, and obtain the query text carried in the search instruction;

[0007] Call the fine-tuned large model to perform text conversion on the query text, convert the non-standard words in the query text into standard words, and obtain the converted target query text;

[0008] Search for the target text matching the target query text in the database and output it.

[0009] Further, the construction process of the database includes:

[0010] Obtain each original text to be stored from a preset data source;

[0011] For each original text, call the large model to correct the original text respectively, convert the original text into a standard text carrying standard words; store the original text and the standard text correspondingly in the database.

[0012] Further, searching for the target text matching the target query text in the database includes:

[0013] For each standard vocabulary carried in the target query text, search in the database for the first candidate standard text containing each standard vocabulary;

[0014] Determine the first original text corresponding to the first candidate standard text in the database as the target text.

[0015] Furthermore, the database also stores the first feature vector corresponding to each standard text;

[0016] The process of searching for the target text matching the target query text in the database includes:

[0017] Determine the second feature vector corresponding to the target query text;

[0018] Search in the database for the target first feature vector matching the second feature vector;

[0019] Determine the second original text corresponding to the second candidate standard text of the target first feature vector in the database as the target text.

[0020] Furthermore, the fine-tuning process of the large model includes:

[0021] Obtain a pre-configured training data set, where the training data set contains each sample text corresponding to each scenario in each field, and each sample text contains at least one standard vocabulary;

[0022] Use the LoRA fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0023] Furthermore, the method further includes:

[0024] Obtain a pre-configured test data set, where the training data set contains each test text corresponding to each scenario in each field;

[0025] Use a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain the evaluation index corresponding to the fine-tuned large model;

[0026] If the evaluation index does not meet the preset requirements, fine-tune the large model again.

[0027] In a second aspect, an embodiment of the present application further provides a search device, and the device includes:

[0028] A processing module, configured to receive an input search instruction, obtain the query text carried in the search instruction; call the fine-tuned large model to perform text conversion on the query text, convert non-standard words in the query text into standard words, and obtain the converted target query text.

[0029] A search module, configured to search for and output the target text in the database that matches the target query text.

[0030] Further, the processing module is further configured to obtain each original text to be stored from a preset data source; for each original text, call the large model to respectively correct the original text, convert the original text into a standard text carrying standard words; and correspondingly save the original text and the standard text into the database.

[0031] In a third aspect, an embodiment of the present application further provides an electronic device, which at least includes a processor and a memory. When the processor executes a computer program stored in the memory, the steps of the search method as described in any one of the above are implemented.

[0032] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the search method as described in any one of the above are implemented.

[0033] In an embodiment of the present application, an input search instruction is received, the query text carried in the search instruction is obtained; the fine-tuned large model is called to perform text conversion on the query text, non-standard words in the query text are converted into standard words, and the converted target query text is obtained; the target text in the database that matches the target query text is searched for and output. In an embodiment of the present application, the input query text is polished through the fine-tuned large model to obtain the target query text carrying standard words, thereby enhancing the matching degree between the target query text and the text stored in the database. Not only can the required target text be found faster, but also the quality of the search results is significantly improved. Description of the Drawings

[0034] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 A schematic diagram of a search process provided by an embodiment of the present application;

[0036] Figure 2The search flow chart provided by the embodiments of the present application;

[0037] Figure 3 The schematic structural diagram of a search device provided by the embodiments of the present application;

[0038] Figure 4 The schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0039] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application. Without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0040] The terms "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application may represent at least two, for example, it may be two, three or more, and the embodiments of the present application do not make limitations.

[0041] The following makes an explanation of the exemplary embodiments of the present application with reference to the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, and they should be considered only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope of the disclosure of the present application. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures. It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0042] With the development of Internet technology, search engines have become an important tool for people to obtain information. However, traditional search engines mainly rely on word matching based on word segmentation and semantic matching technology based on text embedding. When facing the language differences between users and documents, these methods often fail to achieve ideal recall effects.

[0043] During the search process, the electronic device can match the text through word matching based on word segmentation and also through semantic matching based on text embedding. Among them, word matching based on word segmentation decomposes the text into independent lexical units and realizes the matching by comparing the query words with the words in the document. Although it is simple and efficient, it has limitations in dealing with synonyms, near-synonyms and context understanding. Semantic matching based on text embedding uses deep learning technology to convert the text into vector form and measures the semantic relevance of the text by calculating the similarity between vectors. Although it can better capture the deep meaning of the text, it still needs to be improved for the diversity and complexity of language expressions.

[0044] Based on this, in order to enhance the matching degree between the user query and the text stored in the database and improve the search quality, the embodiments of the present application provide a search method, device, equipment and medium.

[0045] In the embodiments of the present application, receive the input search instruction, and obtain the query text carried in the search instruction; call the fine-tuned large model to perform text conversion on the query text, convert the non-standard words in the query text into standard words, and obtain the converted target query text; search for the target text in the database that matches the target query text and output it.

[0046] Embodiment 1:

[0047] Figure 1 FIG. is a schematic diagram of a search process provided by the embodiments of the present application, and the process includes:

[0048] S101: Receive the input search instruction, and obtain the query text carried in the search instruction.

[0049] A search method provided by the embodiments of the present application is applied to an electronic device, which can be a PC, a server, etc., and a search engine is deployed in the electronic device.

[0050] In the embodiments of the present application, the user can search for relevant content through the search engine deployed in the electronic device, including but not limited to noun explanations, knowledge Q&A, etc. The user can input a search instruction carrying the query text to the electronic device, and after the electronic device receives the search instruction, it responds to the search instruction.

[0051] Specifically, in the embodiments of the present application, a user can input a search instruction to an electronic device in text form through external devices such as a keyboard and a mouse. Alternatively, an audio acquisition component of the electronic device collects the user's voice, converts the voice into text, and uses the converted text as a search instruction.

[0052] After the electronic device receives the search instruction input by the user, the electronic device obtains the query text carried in the search instruction and searches for content matching the query text in the database.

[0053] S102: Invoke the fine-tuned large model to perform text conversion on the query text, convert non-standard words in the query text into standard words, and obtain the converted target query text.

[0054] In a possible implementation process, since the search instruction is input by the user, and the user is very likely to be a non-technical person, the input text may contain non-standard words, etc. If the electronic device directly searches based on the query text carried in the user input search instruction, it is very likely that the content searched is not what the user needs. Among them, non-standard words are non-professional terms, colloquial words, such as "driving after drinking" and so on.

[0055] Based on this, in order to improve the accuracy and usability of the search, before the search, the electronic device can first polish the query text, convert non-standard words in the query text into standard words, obtain the target query text, and then search based on the target query text.

[0056] Specifically, a fine-tuned large model for polishing the query text is configured in the electronic device. The electronic device inputs the query text to be polished into the large model. The large model will identify the non-standard words contained in the query text and convert the non-standard words into standard words to obtain the converted query text. The large model uses the converted query text as the target query text and outputs it.

[0057] For example, in the actual application process, if the query text contains the non-standard word "driving after drinking", after the large model receives the query text, it will convert "driving after drinking" into "drunk driving" and output the converted target query text.

[0058] It should be noted that a query text may contain one non-standard word or multiple non-standard words, which is not limited here. Regardless of how many non-standard words are contained in the query text, the large model will convert all the non-standard words contained in the query text into the corresponding standard words and then output the converted target query text.

[0059] In addition, the electronic device can also use a large model to polish the query text in multiple styles to obtain multiple target query texts, and then perform a search based on the multiple target query texts, which helps to bridge the language gap between the user and the document.

[0060] S103: Find the target text in the database that matches the target query text and output it.

[0061] In the embodiment of the present application, after the electronic device receives the target query text output by the large model, the electronic device searches in the database for the target text that matches the target query text according to the target query text, and the electronic device outputs the target text.

[0062] If the electronic device determines multiple target query texts, the electronic device determines the target text respectively matched by each target query text, and then the electronic device outputs each target text for the user to freely select.

[0063] Specifically, in different style dimensions, the query content input by the user and the database will perform matching and recall operations by themselves. Through multi-angle matching, the relevance and accuracy of the recall results are improved.

[0064] In the embodiment of the present application, receive the input search instruction, obtain the query text carried in the search instruction; call the fine-tuned large model to perform text conversion on the query text, convert the non-standard vocabulary in the query text into standard vocabulary, and obtain the converted target query text; find the target text in the database that matches the target query text and output it. In the embodiment of the present application, the input query text is polished by the fine-tuned large model to obtain the target query text carrying standard vocabulary, thereby enhancing the matching degree between the target query text and the text stored in the database, not only being able to find the required target text faster, but also significantly improving the quality of the search results.

[0065] Embodiment 2:

[0066] In order to enhance the matching degree between the user query and the text stored in the database and improve the search quality, on the basis of the above embodiment, in the embodiment of the present application, the construction process of the database includes:

[0067] Obtain each original text to be stored from the preset data source;

[0068] For each original text, call the large model to respectively correct the original text, convert the original text into a standard text carrying standard vocabulary; save the original text and the standard text correspondingly to the database.

[0069] In the embodiments of the present application, the database can be understood as an ES document collection. The electronic device can obtain texts in multiple fields and form an ES document collection with each obtained text. The electronic device uses this ES document collection as the database.

[0070] In a possible implementation manner, after the electronic device obtains each text, for each text, the text can be used as the original text, and the fine-tuned large model can be used to polish the original text to achieve the effect of correcting the original text.

[0071] Specifically, in the embodiments of the present application, the electronic device obtains each original text to be stored from a preset data source. For each original text, the electronic device calls the large model to correct the original text respectively, converts the original text into a standard text carrying standard vocabulary; and stores the original text and the standard text correspondingly in the database.

[0072] In addition, to increase the diversity of text expressions, in the embodiments of the present application, when the electronic device calls the large model to polish the original text, it can choose to polish the original text in multiple styles to obtain standard texts in different styles. The electronic device stores the original text and the corresponding standard texts in various styles in the database.

[0073] Embodiment 3:

[0074] To enhance the matching degree between the user query and the texts stored in the database and improve the search quality, based on the above embodiments, in the embodiments of the present application, finding the target text in the database that matches the target query text includes:

[0075] According to each standard vocabulary carried in the target query text, find the first candidate standard text in the database that contains each standard vocabulary;

[0076] Determine the first original text corresponding to the first candidate standard text in the database as the target text.

[0077] In the embodiments of the present application, when the electronic device searches for the target text in the database that matches the target query text, it can query based on the keyword matching method.

[0078] Specifically, the electronic device obtains each standard vocabulary carried in the target query text and queries the first candidate standard text in the database that contains each standard vocabulary. The electronic device determines the original text corresponding to the first candidate standard text as the target text.

[0079] It should be noted that in the embodiments of the present application, when the electronic device performs matching, all the standard vocabulary carried in the target query text is included in the first candidate standard text obtained by matching. Moreover, the number of the first candidate standard texts obtained by matching may be one or multiple. If the number of the first candidate standard texts is multiple, the original text corresponding to each first candidate standard text is respectively determined as the target text.

[0080] In addition, it may also occur that two or more first candidate standard texts correspond to the same original text, which is caused by generating standard texts in multiple styles. For example, the standard vocabulary corresponding to the non-standard vocabulary "driving after drinking alcohol" can be "drunk driving", and can also be "driving under the influence".

[0081] Embodiment 4:

[0082] In order to enhance the matching degree between the user query and the text stored in the database and improve the search quality, on the basis of the above embodiments, in the embodiments of the present application, the database further stores a first feature vector corresponding to each standard text;

[0083] The process of finding the target text in the database that matches the target query text includes:

[0084] Determine a second feature vector corresponding to the target query text;

[0085] Find a target first feature vector in the database that matches the second feature vector;

[0086] Determine the second original text corresponding to the second candidate standard text of the target first feature vector in the database as the target text.

[0087] In the embodiments of the present application, the electronic device may also determine the target text in the database that matches the target query text based on a similarity algorithm.

[0088] In a possible implementation manner, the electronic device uses a preset text embedding algorithm to determine a first feature vector corresponding to each standard text in the database, and then determines a second feature vector corresponding to the target query text. The electronic device searches for a target first feature vector that matches the second feature vector among each first feature vector according to the second feature vector.

[0089] Specifically, the electronic device may determine the similarity between each first feature vector and the second feature vector according to a preset similarity algorithm, and the electronic device determines the first feature vector corresponding to the highest similarity as the target first feature vector. Among them, the preset similarity algorithm may be a cosine similarity algorithm, or an Euclidean distance similarity algorithm, etc., which is not limited herein.

[0090] In addition, in order to reduce the load of the electronic device and improve the search efficiency at the same time, in the embodiments of the present application, the electronic device can pre-determine the first feature vector corresponding to each standard text, and store each standard text and the first feature vector in the database in a corresponding manner. Based on this, during the search process, the electronic device can directly obtain each first feature vector from the database without further calculation, thereby reducing the workload of the electronic device and improving the search efficiency.

[0091] Embodiment 5:

[0092] In order to enhance the matching degree between the user query and the text stored in the database and improve the search quality, on the basis of the above embodiments, in the embodiments of the present application, the fine-tuning process of the large model includes:

[0093] Obtain a pre-configured training data set, where each sample text corresponding to each scenario in each field is included in the training data set, and at least one standard vocabulary is included in each sample text;

[0094] Use the LoRA fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0095] In the embodiments of the present application, in order to enable the large model to be used for text polishing, the electronic device will use a large amount of data to fine-tune the original large model, so that the large model learns the text polishing ability.

[0096] Specifically, a training data set is pre-configured in the electronic device. The training data set includes high-quality sample texts in multiple fields and multiple scenarios. The electronic device uses the LoRA fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0097] It should be noted that, in the embodiments of the present application, in order to ensure that the large model can understand and generate high-quality text in specific fields such as law, education, and medical care, in the embodiments of the present application, detailed data preparation steps are designed for the texts in these three fields, including but not limited to the following:

[0098] I. Legal field:

[0099] Data collection: Collect a large amount of text from official judicial document databases, legal literature, regulations, and case analyses written by lawyers and judges.

[0100] Data cleaning: Remove duplicates, informal language expressions, and standardize the documents, such as unifying the use of terms.

[0101] Data annotation: Invite professional legal personnel to participate and annotate key concepts, terms, and legal relationships so that the large model can learn the correct legal logic.

[0102] II. Education field:

[0103] Data collection: Cover educational resources such as textbooks, academic papers, online course materials, and teacher lecture notes at all levels from primary school to university.

[0104] Data cleaning: Eliminate irrelevant information (such as advertisements) to ensure the educational and scientific nature of the content; at the same time, adjust the text format to make it more machine-readable.

[0105] Data augmentation: Increase interactivity in the form of question-and-answer pairs, simulate the scenarios of students asking questions and teachers answering, and strengthen the support for conversational learning.

[0106] III. Medical field:

[0107] Data collection: Integrate data from multiple sources such as clinical guidelines, medical journal articles, patient medical record summaries, and health popular science articles.

[0108] Data privacy protection: Strictly comply with regulations to ensure that all personal sensitive information is anonymized or de-sensitized.

[0109] Professional knowledge correction: Reviewed and corrected by medical experts to ensure the accuracy and reliability of the information, with special attention to content such as disease descriptions, treatment plans, and drug instructions.

[0110] In the embodiments of the present application, after constructing the training data set, the electronic device can fine-tune the original large model, including but not limited to the following processes:

[0111] 1. Model selection: Select an advanced pre-trained language model as the basic architecture, such as BERT or GPT series, which have demonstrated excellent performance in general language tasks.

[0112] 2. Domain adaptation: Based on the above-selected basic model, perform domain-specific fine-tuning. This includes but not limited to adjusting the model structure and optimizing hyperparameter settings to better fit the characteristics of the target domain.

[0113] 3. Cultivation of multi-style generation ability: Introduce an adversarial training mechanism or adopt other strategies to enable the model to learn to flexibly switch the writing style according to different scenarios and user needs. For example, maintain formality and rigor during legal consultation, and be more friendly and easy to understand during educational assistance.

[0114] 4. Continuous iterative improvement: As more high-quality domain data accumulates and technology develops, regularly update the model weights to maintain its cutting-edge nature and practicality. In addition, a feedback loop system should be established to guide the development direction of subsequent versions based on user evaluations in actual applications.

[0115] Example 6:

[0116] To enhance the matching degree between user queries and the text stored in the database and improve the search quality, based on the above embodiments, in the embodiments of the present application, the method further includes:

[0117] Obtain a pre-configured test data set, where each test text corresponding to each scenario in each domain is included in the training data set;

[0118] Use a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain the evaluation indicators corresponding to the fine-tuned large model;

[0119] If the evaluation indicators do not meet the preset requirements, fine-tune the large model again.

[0120] In the embodiments of the present application, after the large model is fine-tuned, the electronic device will also test the text polishing ability of the large model. If the test is unqualified, the large model needs to be fine-tuned again. If the test is qualified, the large model can be put into use.

[0121] In a possible implementation manner, a test data set is pre-configured in the electronic device. The electronic device uses a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain the evaluation indicators corresponding to the fine-tuned large model; if the evaluation indicators do not meet the preset requirements, fine-tune the large model again.

[0122] Specifically, in the embodiments of the present application, diverse evaluation indicators such as BLEU score, ROUGE-L, METEOR, etc. can be used as automatic evaluation indicators, combined with manual review, to ensure that the model output not only meets professional standards but also is close to human intuition.

[0123] Figure 2 The search flow chart provided for the embodiments of the present application is as Figure 2 shown, and this process includes:

[0124] 1. Large model training: Based on the prepared training data set, construct and train a powerful large model for multi-style text generation and understanding.

[0125] 2. Document Polishing: Using the trained large model, polish the documents stored in Elasticsearch (ES) in multiple styles to generate N versions of the document content in different styles, increasing the diversity of document expression and making it more likely to cover the user's query intent.

[0126] 3. Query Polishing: When the user submits a query request, also use the large model to polish the query text in multiple styles to generate N versions of the target query text, which helps to bridge the language gap between the user and the document.

[0127] 4. Matching and Retrieval: On different style dimensions, perform matching and retrieval operations on the query content input by the user and the ES document set respectively. Through multi-angle matching, the relevance and accuracy of the retrieval results are improved.

[0128] Through practical application verification, the embodiments of this application have significantly improved the performance of the search system, especially in professional fields such as law, education, and medicine, and received good user feedback. Users can not only find the required information faster, but also the quality of the search results has been significantly improved.

[0129] In summary, the search method based on large model polishing provided by the embodiments of this application provides a new idea for solving the problems existing in traditional search technologies. By enhancing the matching degree between the user query and the document content, the search experience is effectively improved, providing users with more accurate and personalized information services. With the continuous progress of NLP technology, it is expected that this solution will be widely applied in more fields, further promoting the development of information retrieval technology.

[0130] Embodiment 7:

[0131] Based on the above embodiments, the embodiments of this application also provide a search device. Figure 3 The following is a schematic structural diagram of a search device provided by the embodiments of this application. The device includes:

[0132] A processing module 301, configured to receive an input search instruction, obtain the query text carried in the search instruction; call the fine-tuned large model to perform text conversion on the query text, and convert the non-standard words in the query text into standard words to obtain the converted target query text;

[0133] A search module 302, configured to search for and output the target text matching the target query text in the database.

[0134] In a possible implementation, the processing module 301 is further configured to obtain each original text to be stored from a preset data source; for each original text, call the large model to correct the original text respectively, and convert the original text into a standard text carrying standard vocabulary; save the original text and the standard text correspondingly to the database.

[0135] In a possible implementation, the search module 302 is specifically configured to, according to each standard vocabulary carried in the target query text, search in the database for a first candidate standard text containing each standard vocabulary; determine the first original text corresponding to the first candidate standard text in the database as the target text.

[0136] In a possible implementation, the database also stores a first feature vector corresponding to each standard text;

[0137] The search module 302 is specifically configured to determine a second feature vector corresponding to the target query text; search in the database for a target first feature vector that matches the second feature vector; determine the second original text corresponding to the second candidate standard text of the target first feature vector in the database as the target text.

[0138] In a possible implementation, the processing module 301 is further configured to obtain a pre-configured training data set, where the training data set contains each sample text corresponding to each scenario in each field, and each sample text contains at least one standard vocabulary; use the LoRA fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0139] In a possible implementation, the processing module 301 is further configured to obtain a pre-configured test data set, where the training data set contains each test text corresponding to each scenario in each field; use a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain an evaluation index corresponding to the fine-tuned large model; if the evaluation index does not meet the preset requirements, fine-tune the large model again.

[0140] Example 8:

[0141] Based on the above embodiments, an embodiment of the present application further provides an electronic device, Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present application is shown in Figure 4As shown in the figure, it includes: a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404;

[0142] The memory 403 stores a computer program. When the program is executed by the processor 401, the processor 401 is caused to execute the following steps:

[0143] Receive an input search instruction and obtain the query text carried in the search instruction;

[0144] Call the fine-tuned large model to perform text conversion on the query text, convert the non-standard vocabulary in the query text into standard vocabulary, and obtain the converted target query text;

[0145] Search for the target text in the database that matches the target query text and output it.

[0146] In a possible implementation manner, the construction process of the database includes:

[0147] Obtain each original text to be stored from a preset data source;

[0148] For each original text, call the large model to respectively correct the original text, convert the original text into a standard text carrying standard vocabulary; and store the original text and the standard text correspondingly in the database.

[0149] In a possible implementation manner, the searching for the target text in the database that matches the target query text includes:

[0150] According to each standard vocabulary carried in the target query text, search in the database for the first candidate standard text containing each standard vocabulary;

[0151] Determine the first original text corresponding to the first candidate standard text in the database as the target text.

[0152] In a possible implementation manner, the database also stores the first feature vector corresponding to each standard text;

[0153] The searching for the target text in the database that matches the target query text includes:

[0154] Determine the second feature vector corresponding to the target query text;

[0155] Search in the database for the target first feature vector that matches the second feature vector;

[0156] Determine the second original text corresponding to the second candidate standard text of the target first feature vector in the database as the target text.

[0157] In a possible implementation manner, the fine-tuning process of the large model includes:

[0158] Obtain a pre-configured training data set, where the training data set contains each sample text corresponding to each scenario in each field, and each sample text contains at least one standard vocabulary;

[0159] Use the LoRA fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0160] In a possible implementation manner, the method further includes:

[0161] Obtain a pre-configured test data set, where the training data set contains each test text corresponding to each scenario in each field;

[0162] Use a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain an evaluation index corresponding to the fine-tuned large model;

[0163] If the evaluation index does not meet the preset requirements, fine-tune the large model again.

[0164] Since the principle of the above electronic device for solving problems is similar to that of the search method, the implementation of the above electronic device can refer to the embodiments of the method, and the repeated parts will not be described again.

[0165] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 402 is used for communication between the above electronic device and other devices. The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0166] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0167] Embodiment 9:

[0168] Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the following steps are implemented when the processor executes:

[0169] Receive an input search instruction, and obtain the query text carried in the search instruction;

[0170] Call the fine-tuned large model to perform text conversion on the query text, convert non-standard words in the query text into standard words, and obtain the converted target query text;

[0171] Search for and output the target text in the database that matches the target query text.

[0172] In a possible implementation manner, the construction process of the database includes:

[0173] Obtain each original text to be stored from a preset data source;

[0174] For each original text, call the large model to correct the original text respectively, convert the original text into a standard text carrying standard words; store the original text and the standard text correspondingly in the database.

[0175] In a possible implementation manner, the searching for the target text in the database that matches the target query text includes:

[0176] According to each standard word carried in the target query text, search in the database for the first candidate standard text containing each standard word;

[0177] Determine the first original text corresponding to the first candidate standard text in the database as the target text.

[0178] In a possible implementation manner, the database also stores the first feature vector corresponding to each standard text;

[0179] The searching for the target text in the database that matches the target query text includes:

[0180] Determine the second feature vector corresponding to the target query text;

[0181] Search in the database for the target first feature vector that matches the second feature vector;

[0182] Determine the second original text corresponding to the second candidate standard text of the target first feature vector in the database as the target text.

[0183] In a possible implementation manner, the fine-tuning process of the large model includes:

[0184] Obtain a pre-configured training data set, where each sample text corresponding to each scenario in each field is included in the training data set, and at least one standard vocabulary is included in each sample text;

[0185] Use the lora fine-tuning algorithm and the training data set to fine-tune the original large model to obtain a fine-tuned large model.

[0186] In a possible implementation manner, the method further includes:

[0187] Obtain a pre-configured test data set, where each test text corresponding to each scenario in each field is included in the training data set;

[0188] Use a preset instruction evaluation algorithm and the test data set to test the fine-tuned large model to obtain the evaluation index corresponding to the fine-tuned large model;

[0189] If the evaluation index does not meet the preset requirements, fine-tune the large model again.

[0190] Since the principle of the above computer program product for solving problems is similar to the search method, the implementation of the above computer program product can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0191] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0192] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to this application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0193] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0194] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0195] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A search method, characterized in that: The method comprises: Receiving an input search instruction and obtaining a query text carried in the search instruction; Calling the fine-tuned large model to perform text conversion on the query text, converting non-standard words in the query text into standard words, and obtaining a converted target query text; Search the database for target text that matches the target query text and output it.

2. The method according to claim 1, characterized in that The database construction process includes: Obtain each original text to be stored from a preset data source; For each original text, the large model is called to modify the original text respectively, and the original text is converted into a standard text carrying a standard vocabulary; the original text and the standard text are correspondingly saved in the database.

3. The method according to claim 2, characterized in that The target text in the search database that matches the target query text includes: According to each standard word carried in the target query text, searching the database for a first candidate standard text containing each standard word; A first original text corresponding to the first candidate standard text in the database is determined as the target text.

4. The method according to claim 2, characterized in that: The database also stores a first feature vector corresponding to each standard text; The target text in the search database that matches the target query text includes: Determine a second feature vector corresponding to the target query text; Searching the database for a target first feature vector that matches the second feature vector; A second original text corresponding to the second candidate standard text of the target first feature vector in the database is determined as the target text.

5. The method according to claim 1, characterized in that The fine-tuning process of the large model includes: Obtain a preconfigured training data set, wherein the training data set includes each sample text corresponding to each scene in each field, and each sample text includes at least one standard vocabulary; The lora fine-tuning algorithm and the training data set are used to fine-tune the original large model to obtain a fine-tuned large model.

6. The method according to claim 5, characterized in that The method further comprises: Obtain a preconfigured test data set, wherein the training data set contains each test text corresponding to each scene in each field; Using a preset instruction evaluation algorithm and the test data set, the fine-tuned large model is tested to obtain an evaluation index corresponding to the fine-tuned large model; If the evaluation index does not meet the preset requirements, the large model is fine-tuned again.

7. A search device, characterized in that: The device comprises: A processing module is used to receive an input search instruction and obtain a query text carried in the search instruction; call the fine-tuned large model to perform text conversion on the query text, convert non-standard words in the query text into standard words, and obtain a converted target query text; The search module is used to search for target text matching the target query text in a database and output it.

8. The device according to claim 7, characterized in that The processing module is also used to obtain each original text to be stored from a preset data source; for each original text, call the large model to modify the original text respectively, and convert the original text into a standard text carrying a standard vocabulary; and save the original text and the standard text correspondingly in the database.

9. An electronic device, characterized in that: The electronic device comprises at least a processor and a memory, and the processor is used to implement the steps of the search method according to any one of claims 1 to 6 when executing a computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the steps of the search method according to any one of claims 1 to 6.