Information retrieval method and device, electronic equipment and readable storage medium

By receiving user query information, generating reference answers and performing mixed searches, and generating target answers using large language models, the problem of insufficient relevance and accuracy in traditional systems is solved, and more efficient information retrieval is achieved.

CN120336487AInactive Publication Date: 2025-07-18SHENZHEN ZHUOYUE ZHIYUN TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510484402.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional search-based systems have problems of insufficient correlation and accuracy when facing large-scale corpus, especially when dealing with complex problems, the accuracy of search results is low.

Method used

By receiving the query information input by the user, the query database is determined, and reference answers are generated based on the local model, the query information and reference answers are spliced and mixed searches are performed, and the target answers are generated using the preset large language model, and the results are displayed through digital people.

Benefits of technology

It improves the relevance and accuracy of the search results, can understand the query intention more accurately, and improves the effectiveness of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336487A_ABST
    Figure CN120336487A_ABST
Patent Text Reader

Abstract

The invention discloses an information retrieval method and device, electronic equipment and a readable storage medium, and the information retrieval method comprises the following steps: receiving query information input by a user, and determining a query database corresponding to the query information; generating a reference answer corresponding to the query information based on the query database and a local model; splicing the query information and the reference answer to obtain updated query information; performing mixed retrieval on the updated query information, and processing a document obtained by the mixed retrieval; and based on a preset cue word and the processed document, generating a target answer corresponding to the query information through a preset large language model, and displaying the target answer through a digital person. According to the information retrieval scheme provided by the invention, the correlation and the accuracy of retrieval results can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and specifically relates to an information retrieval method, device, electronic device, and readable storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, knowledge retrieval and generation systems have been widely used in various industries, especially in the fields of enterprise knowledge bases, intelligent assistants, virtual customer service, etc. These systems improve the efficiency and accuracy of information acquisition by quickly retrieving and generating relevant answers from massive amounts of information.

[0003] However, with the continuous increase of large-scale corpora, traditional retrieval-based systems (such as the BM25-based retrieval model) often face problems of insufficient relevance and accuracy. These systems usually retrieve relevant information based on keyword matching or simple similarity calculation, but due to the lack of in-depth semantic understanding, they are prone to returning answers that do not exactly match the query intent. Especially when dealing with complex problems, the accuracy of the retrieval results is relatively low. Summary of the Invention

[0004] In view of the above technical problems, this application provides an information retrieval method, device, electronic device, and readable storage medium, which can improve the accuracy of retrieval results.

[0005] To solve the above technical problems, this application provides an information retrieval method, including:

[0006] Receiving query information input by a user and determining a query database corresponding to the query information;

[0007] Generating a reference answer corresponding to the query information based on the query database and a local model;

[0008] Concatenating the query information and the reference answer to obtain updated query information;

[0009] Performing hybrid retrieval on the updated query information and processing the documents obtained by the hybrid retrieval;

[0010] Generating a target answer corresponding to the query information based on a preset prompt word and the processed documents through a preset large language model, and presenting the target answer through a digital human.

[0011] Optionally, in some embodiments of this application, the performing hybrid retrieval on the updated query information and processing the documents obtained by the hybrid retrieval includes:

[0012] Extracting a query vector corresponding to the updated query information;

[0013] Retrieve the updated query information based on the query vector, and;

[0014] Perform keyword retrieval on the updated query information;

[0015] Fuse the vector retrieval result and the keyword retrieval result.

[0016] Optionally, in some embodiments of the present application, the retrieving the updated query information based on the query vector includes:

[0017] Input the query vector into a preset vector database, and calculate the cosine similarity between the query vector and each document vector in the vector database;

[0018] Add the documents corresponding to the document vectors with cosine similarity greater than the preset value to the vector retrieval result.

[0019] Optionally, in some embodiments of the present application, the performing keyword retrieval on the updated query information includes:

[0020] Extract keywords from the updated query information;

[0021] Calculate the term weights and term frequencies of the extracted keywords in each document of the target document library;

[0022] Based on the term weights and terms, and the document lengths of each document in the target document library, calculate the relevance between each document in the target document library and the updated query information;

[0023] Based on the relevance, determine the documents associated with the updated query information in the target document library.

[0024] Optionally, in some embodiments of the present application, it further includes:

[0025] Perform chunking processing on the documents in the target document library;

[0026] Generate embedding vectors for each document chunk through a pre-trained language model;

[0027] Transmit the generated document chunks and embedding vectors to the edge device through the network, so that after receiving the data, the edge device stores the generated document chunks and embedding vectors in the local cache according to the set caching policy.

[0028] Optionally, in some embodiments of the present application, generating the target answer corresponding to the query information through a preset large language model based on the preset prompt words and the processed documents, and displaying the target answer through a digital human, includes:

[0029] Combine the query information, the preset prompt words, and the processed document to obtain an input set;

[0030] Input the input set into a preset large language model to generate a target answer that meets the prompt words;

[0031] Generate an interactive special effect for the target answer, and display the target answer through a digital human according to the interactive special effect.

[0032] Optionally, in some embodiments of the present application, determining the query database corresponding to the query information includes:

[0033] Identify the keywords in the query information;

[0034] Based on the string matching strategy and the keywords, determine the query database corresponding to the query information.

[0035] Optionally, in some embodiments of the present application, it further includes:

[0036] If no keywords in the query information are identified, generate a query vector corresponding to the query information;

[0037] Obtain the average vector corresponding to each preset database;

[0038] Determine the query database corresponding to the query information according to the similarity between the query vector and each average vector.

[0039] Correspondingly, the present application also provides an information retrieval device, including:

[0040] A receiving module, configured to receive query information input by a user and determine the query database corresponding to the query information;

[0041] A generating module, configured to generate a reference answer corresponding to the query information based on the query database and a local model;

[0042] A splicing module, configured to splice the query information and the reference answer to obtain updated query information;

[0043] A retrieval module, configured to perform hybrid retrieval on the updated query information and process the documents obtained by the hybrid retrieval;

[0044] A display module, configured to generate a target answer corresponding to the query information through a preset large language model based on the preset prompt words and the processed documents, and display the target answer through a digital human.

[0045] The present application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0046] The present application also provides a computer storage medium. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0047] As described above, the present application provides an information retrieval method, device, electronic device and readable storage medium. After receiving the query information input by the user and determining the query database corresponding to the query information, a reference answer corresponding to the query information is generated based on the query database and the local model. Then, the query information and the reference answer are spliced to obtain the updated query information. Then, a hybrid retrieval is performed on the updated query information, and the documents obtained by the hybrid retrieval are processed. Finally, based on the preset prompt words and the processed documents, a target answer corresponding to the query information is generated through a preset large language model, and the target answer is displayed through a digital human. In the information retrieval solution provided by the present application, a reference answer corresponding to the query information can be generated based on the query database and the local model. On this basis, the query information and the reference answer are spliced, and a hybrid retrieval is performed on the updated query information obtained by splicing, so that the query intention can be understood more accurately, thereby improving the relevance and accuracy of the retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 is a schematic structural diagram of the information retrieval system provided by the embodiment of the present application;

[0050] Figure 2 is a schematic flowchart of the information retrieval method provided by the embodiment of the present application;

[0051] Figure 3 is a generation system of artificial intelligence knowledge retrieval based on digital human real-time interaction and edge device deployment provided by the embodiment of the present application;

[0052] Figure 4 is a schematic flowchart of the artificial intelligence knowledge retrieval based on digital human real-time interaction and edge device deployment provided by the embodiment of the present application;

[0053] Figure 5 It is a schematic structural diagram of an information retrieval device provided by an embodiment of the present application;

[0054] Figure 6 It is a schematic structural diagram of an intelligent terminal provided by an embodiment of the present application.

[0055] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Embodiments

[0056] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0057] It should be noted that in this document, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanations in the specific embodiments or further in combination with the context of the specific embodiments.

[0058] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] In subsequent descriptions, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the convenience of the description of the present application, and it has no specific meaning in itself. Therefore, "module", "component" or "unit" can be used interchangeably.

[0060] The following specifically describes the embodiments involved in this application. It should be noted that the description order of the embodiments in this application does not limit the priority order of the embodiments.

[0061] The embodiments of this application provide an information retrieval method, device, storage medium, and electronic device. Specifically, the information retrieval method of the embodiments of this application can be executed by an electronic device or a server. Among them, the electronic device can be a terminal. The terminal can be an electronic device such as a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), etc. The terminal can also include a client, and the client can be a media playback client or an instant messaging client, etc.

[0062] For example, when the information retrieval method runs on an electronic device, the electronic device can receive the query information input by the user, determine the query database corresponding to the query information, and then, based on the query database and the local model, the electronic device generates a reference answer corresponding to the query information. Then, the electronic device splices the query information and the reference answer to obtain the updated query information. Then, the electronic device performs a hybrid retrieval on the updated query information and processes the documents obtained from the hybrid retrieval. Finally, based on the preset prompt words and the processed documents, the electronic device generates the target answer corresponding to the query information through the preset large language model and displays the target answer through a digital human. The electronic device can interact with the user through a graphical user interface. The way the electronic device provides the graphical user interface to the user can include various methods. For example, it can be rendered and displayed on the display screen of the electronic device, or the graphical user interface can be presented through holographic projection. For example, the electronic device can include a touch display screen and a processor, and the touch display screen is used to present the graphical user interface and receive the operation instructions generated by the user acting on the graphical user interface.

[0063] Please refer to Figure 1 , Figure 1System schematic diagram of the information retrieval device provided by the embodiment of the present application. The system may include at least one electronic device 1000 and at least one server or personal computer 2000. The electronic device 1000 held by the user can be connected to different servers or personal computers through a network. The electronic device 1000 can be an electronic device with computing hardware that can support and execute software products corresponding to multimedia. In addition, the electronic device 1000 can also have one or more multi-touch sensitive screens for sensing and obtaining inputs of touch or swipe operations performed by the user at multiple points on one or more touch display screens. In addition, the electronic device 1000 can be interconnected with the server or personal computer 2000 through a network. The network can be a wireless network or a wired network. For example, the wireless network can be a wireless local area network (WLAN), local area network (LAN), cellular network, 2G network, 3G network, 4G network, 5G network, etc. In addition, different electronic devices 1000 can also use their own Bluetooth network or hotspot network to connect to other embedded platforms or connect to servers and personal computers, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0064] The embodiment of the present application provides an information retrieval method, which can be executed by an electronic device or a server. The embodiment of the present application takes the case where the information retrieval method is executed by an electronic device as an example for illustration. Among them, the electronic device includes a touch display screen and a processor. The touch display screen is used to present a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. When the user operates on the graphical user interface through the touch display screen, the graphical user interface can control the content on the local side of the electronic device in response to the received operation instructions, or can also control the content on the server side in response to the received operation instructions. For example, the operation instructions generated by the user acting on the graphical user interface include instructions for processing initial audio data, and the processor is configured to start the corresponding application program after receiving the instructions provided by the user. In addition, the processor is configured to render and draw a graphical user interface associated with the application program on the touch display screen. The touch display screen is a multi-touch sensitive screen that can sense touch or swipe operations performed simultaneously at multiple points on the screen. When the user performs a touch operation on the graphical user interface with a finger, when the graphical user interface detects the touch operation, it controls the corresponding operation to be displayed in the graphical user interface of the application.

[0065] The information retrieval solution provided by this application can generate a reference answer corresponding to the query information based on the query database and the local model. On this basis, the query information and the reference answer are spliced, and the updated query information obtained by splicing is subjected to hybrid retrieval, so as to more accurately understand the query intention, thereby improving the relevance and accuracy of the retrieval results.

[0066] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0067] An information retrieval method includes: receiving query information input by a user and determining a query database corresponding to the query information; generating a reference answer corresponding to the query information based on the query database and the local model; splicing the query information and the reference answer to obtain updated query information; performing hybrid retrieval on the updated query information and processing the documents obtained by the hybrid retrieval; generating a target answer corresponding to the query information through a preset large language model based on a preset prompt word and the processed documents, and displaying the target answer through a digital human.

[0068] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of the information retrieval method provided by the embodiment of this application. The specific process of this information retrieval method can be as follows:

[0069] 101. Receive query information input by a user and determine a query database corresponding to the query information.

[0070] The query information is the starting point of the system operation, guiding the system to perform subsequent knowledge retrieval and answer generation. It is actively provided by the user and can be input in the form of voice or text. In daily life and work scenarios, it has various forms. In the intelligent customer service scenario, the user may input a text query such as "My mobile phone package traffic has run out. What should I do?" regarding product usage problems; in the financial consultation scenario, the user asks by voice "What are the low-risk financial products recently?"

[0071] The query database is a resource library that stores information related to a specific field and provides data support for the system to retrieve and generate answers. The query database is closely related to different professional fields, and professional knowledge, cases, terms, etc. in this field are stored in the database. For example, the database in the legal field contains various laws and regulations, actual legal cases and analyses; the database in the financial field covers financial product information, market dynamics, investment strategies, etc.

[0072] For example, an information retrieval system constructs an interactive interface in the form of a digital human, supporting two input methods: voice and text. In actual application scenarios, users can either, like on an intelligent legal consultation platform, input "I have a dispute over the contract I signed. What should I do?" via the keyboard; or, in the scenario of an intelligent health assistant, say to the electronic device "I've been having trouble sleeping lately. What can I do about it?". Then, the query information is matched against a predefined query database. By identifying the key information in the query information, the matching of the query database is carried out.

[0073] Optionally, in some embodiments of the present application, the step of "determining the query database corresponding to the query information" may specifically include:

[0074] Identifying the keywords in the query information;

[0075] Based on the string matching strategy and the keywords, determining the query database corresponding to the query information.

[0076] Keywords are keyword terms extracted from the query information input by the user that can reflect the query theme and the field to which it belongs. Keywords are often related to a specific field and are important clues for determining the query database. For example, when the user inputs "I want to consult about legal issues regarding real estate transactions", the system, through natural language processing technology, identifies keywords such as "real estate transactions" and "law". These keywords can not only reflect the user's query theme but also the professional field to which it belongs. Among them, the word "law" clearly points to the legal field, providing key information for subsequent determination of the query database. Based on the identified keywords, the corresponding query database can be quickly located by means of the string matching strategy. For example, when the user inputs "I want to sue the other party for breach of contract. What should I do?", keywords such as "sue" and "breach of contract" indicate that the query involves the legal field, and the database in the legal field can be found accordingly, providing a direction for subsequent retrieval of relevant information such as legal provisions and cases.

[0077] Specifically, the string matching strategy is pre-set. After identifying the keywords, these strategies are used to determine the corresponding query database. When the keyword "law" is identified, according to the predefined domain information, string rules (such as regular expressions) are used for quick matching. If the identifier or description information of the database contains strings related to "law", the query database in the legal field can be directly located. Suppose there are databases in multiple fields such as law, finance, and medicine in the information retrieval system, and each database has a corresponding identifier, such as "law_database", "finance_database", "medical_database". When the keyword "law" appears, the system can accurately find "law_database" by matching the string "law", thereby determining that this database is the query database corresponding to the query information this time, achieving fast and accurate database positioning and providing a data basis for subsequent retrieval and answer generation.

[0078] Optionally, in some embodiments of the present application, it further includes:

[0079] If the keyword in the query information is not recognized, a query vector corresponding to the query information is generated;

[0080] Obtain the average vector corresponding to each preset database;

[0081] Determine the query database corresponding to the query information according to the similarity between the query vector and each average vector.

[0082] In the embodiments of the present application, if no keyword is recognized, the query information is converted into a query vector. This conversion is achieved by means of a specific embedding model (such as OllamaEmbeddings). The embedding model processes the query information, maps the query information in text form into the vector space, and generates a numerical vector representation. For example, when the user inputs "I always feel weak and listless recently", the information retrieval system cannot recognize obvious domain keywords. At this time, the embedding model will convert this sentence into a vector containing multiple numerical values, and this vector contains the semantic features of the query information. Then, calculate the similarity between the query vector and the average vectors of each preset database. Common calculation methods such as cosine similarity are used. By comparing the magnitudes of the similarities, find the average vector of the preset database with the highest similarity to the query vector, and then determine the database corresponding to this average vector as the query database corresponding to the query information. For example, calculate the cosine similarity between the query vector of "feeling weak" and the average vectors of multiple preset databases such as medical, common sense of life, and sports fitness. If the similarity with the average vector of the medical database is the highest, then determine the medical database as the query database corresponding to this query information. In this way, even in the absence of clear keywords, the system can find the database most likely to contain relevant information through vector similarity calculation, providing a data basis for subsequent retrieval and answer generation.

[0083] It should be noted that in some embodiments of the present application, each preset database can be processed in advance to calculate the average vector corresponding to each database. The specific approach is to first convert the documents in the database into vector form (also using the embedding model), and then perform statistical calculations on these vectors to obtain the average vector representing the overall characteristics of the database. For example, for the medical database, convert the documents such as various disease introductions and treatment methods in it into vectors, and obtain an average vector that can summarize the main content of the medical database by calculating the average of these vectors. Each preset database has such an average vector as its "feature identifier".

[0084] 102. Generate a reference answer corresponding to the query information based on the query database and the local model.

[0085] Specifically, according to the determined query database, the local caching mechanism of the edge device can be used to pre-load some document blocks and embedding vectors in advance. For example, if the query information is "I want to sue the other party for breach of contract. What should I do" and the query database is related to the legal field, relevant document blocks related to legal knowledge will be quickly retrieved from the local cache. If the required information is not in the cache, it will be loaded in real time from the database. These document blocks are the basic data sources for generating the reference answer, and they contain a large amount of information such as professional knowledge and actual cases.

[0086] Then, the query information is sent to the local model without the assistance of external knowledge. The local model can be a lightweight large language model, which, based on its own training data and algorithms, understands and analyzes the query information and attempts to generate a reference answer. For example, when the user queries "What should I do if the other party breaches the contract after signing the contract", the local model may generate a reference answer such as "You can first check the terms regarding breach of contract in the contract and conduct negotiations or take legal measures according to the provisions of the terms". This reference answer contains the local model's preliminary understanding and judgment of the question.

[0087] 103. Concatenate the query information and the reference answer to obtain the updated query information.

[0088] The query information input by the user may have problems such as being vague or brief, and it is difficult to accurately locate relevant documents when used alone for retrieval. The reference answer is generated by the local model in combination with the query database and contains preliminary analysis results and domain-related clues. Concatenating the two can integrate the user's needs and the model's preliminary understanding, provide more comprehensive and accurate semantic information for subsequent retrieval, and improve the relevance and accuracy of the retrieval. Taking the query "investment issues" as an example, the original query is broad. If the reference answer is "You can pay attention to the risks and returns of investment products such as stocks, funds, and bonds", after concatenation, it can clarify the retrieval direction and more accurately locate materials related to investment products.

[0089] Specifically, connect the query information and the reference answer in a specific order to form a new text. For example, place the original query information in the front and the reference answer at the back, and separate them with a specific delimiter (such as a line break, comma, etc.) in the middle to form an updated query information with a clear structure. For example, if the original query information is "How to prevent colds" and the reference answer is "You can prevent colds by strengthening exercise, having a reasonable diet, and getting vaccinated", the concatenation is "How to prevent colds, You can prevent colds by strengthening exercise, having a reasonable diet, and getting vaccinated". In subsequent retrievals, documents can be screened based on both the original requirements and the reference clues.

[0090] 104. Conduct a hybrid retrieval on the updated query information and process the documents obtained from the hybrid retrieval.

[0091] Common retrieval methods include keyword-based retrieval, semantic retrieval, vector retrieval, etc. Keyword-based retrieval finds relevant content by matching the keywords in the query with the words in the text; semantic retrieval focuses on understanding the semantic meaning of the query and can handle language phenomena such as synonyms and near-synonyms; vector retrieval converts the text into a vector representation and conducts retrieval by calculating the similarity between vectors. In actual applications, two or more retrieval methods will be selected for hybridization according to the scenario and requirements. When retrieving highly professional technical documents, a combination of keyword-based retrieval and vector retrieval may be selected.

[0092] In some embodiments of the present application, corresponding weights can be assigned to each retrieval method to determine their importance levels in hybrid retrieval. The weight assignment takes into account multiple factors, such as the accuracy of the retrieval method, the applicable scope, data characteristics, etc. In a dataset containing a large number of professional terms and vocabulary in a specific field, keyword-based retrieval may be assigned a higher weight; while for queries with complex semantics that require context understanding, the weights of semantic retrieval or vector retrieval may be higher.

[0093] After hybrid retrieval, the results obtained by different retrieval methods are merged. Since the results returned by different retrieval methods may overlap or differ, deduplication and integration processing are required to ensure that each result appears only once. According to the pre-assigned weights and the set fusion strategy, the merged results are re-sorted. The fusion strategies include weighted average, voting, etc. For example, for each retrieval result, the corresponding weight can be multiplied by its score in different retrieval methods and then accumulated to obtain the final comprehensive score, and then the results are sorted according to the comprehensive score.

[0094] Specifically, for the documents obtained by hybrid retrieval, a deduplication operation is performed to remove duplicate documents, reduce redundant information, and improve the efficiency of subsequent processing. Whether a document is duplicate can be determined by calculating the feature value of the document (such as a hash value). In addition to sorting in the result fusion stage, the semantic relevance score and context coherence can be further integrated to re-sort the documents. For example, a dedicated re-ranking module, such as LongContextReorder, can be used to make the sorting more reasonable and preferentially return the most relevant documents. In addition, according to requirements, key information can be screened out from the retrieved documents. For example, in a question-and-answer system, paragraphs and sentences directly related to the question are extracted; in an information extraction task, specific types of data fields are extracted, etc. The key information can be determined based on keyword matching, semantic understanding, etc.

[0095] Optionally, in some embodiments of the present application, the step of "performing hybrid retrieval on the updated query information and processing the documents obtained by hybrid retrieval" may specifically include:

[0096] Extracting the query vector corresponding to the updated query information;

[0097] Performing retrieval on the updated query information based on the query vector, and;

[0098] Performing keyword retrieval on the updated query information;

[0099] Fusing the vector retrieval results and the keyword retrieval results.

[0100] For example, specifically, the updated query information can be used to extract the corresponding query vector through an embedding model. Then, the extracted query vector is used to calculate the similarity with the document vectors in the vector database. The higher the similarity, the closer the semantics of the document and the query information. Screening retrieval results: According to the set similarity threshold, the documents corresponding to the document vectors with similarity higher than the threshold are screened out. At the same time, a keyword-based retrieval algorithm such as the BM25 algorithm is used. This algorithm is optimized based on TF-IDF (term frequency-inverse document frequency) weighting and document length normalization to improve the accuracy of keyword matching. Specifically, the search is performed according to the extracted keywords. The algorithm scores and ranks the documents based on factors such as the frequency of keyword occurrence in the document, importance, and document length. Documents with a high frequency of keyword occurrence and rare in other documents will have a higher score and a more forward ranking.

[0101] In addition, weights need to be set for the vector retrieval results and keyword retrieval results respectively to determine their importance in the final retrieval results. By default, the weights of the two can be set to 0.5:0.5, but in actual applications, they will be dynamically adjusted according to the query scenario. According to the set weights, the combined results are re-ranked. For example, for each document, multiply its score in the vector retrieval by the vector retrieval weight, add the score in the keyword retrieval multiplied by the keyword retrieval weight, to obtain a comprehensive score. Then, all documents are sorted according to the comprehensive score to form the final retrieval result.

[0102] Optionally, in some embodiments of the present application, the step of "retrieving the updated query information based on the query vector" may specifically include:

[0103] Input the query vector into a preset vector database, and calculate the cosine similarity between the query vector and each document vector in the vector database;

[0104] Add the documents corresponding to the document vectors with cosine similarity greater than the preset value to the vector retrieval results.

[0105] Before performing the retrieval operation, a large number of documents can be pre-converted into vector form through a local embedding model (such as OllamaEmbeddings) and stored in a preset vector database (such as Chroma vector database). These document vectors represent the semantic features of the documents and constitute a vast semantic vector space.

[0106] During specific retrieval, the query vector corresponding to the updated query information extracted is input into the vector database. This query vector is obtained by processing the updated query information through the same embedding model and contains the semantic information of the user's query. The vector database automatically calculates the cosine similarity between the query vector and each document vector in the database. The calculation is performed using sim(q, d) = q * d / (||q|| * ||d||), where q represents the query vector and d represents the document vector. The value of the cosine similarity ranges from -1 to 1. The closer the value is to 1, the more similar the directions of the two vectors are, which means that the query information represented by the query vector is more semantically similar to the document represented by the document vector; the closer the value is to -1, the more opposite the directions are; when the value is 0, it means that the two vectors are orthogonal, that is, there is no obvious similarity.

[0107] It should be noted that a threshold for the cosine similarity can be preset in advance. The setting of this threshold depends on the specific application scenario and requirements and is an important basis for judging the relevance between the document and the query information. If the threshold is set too high, it may result in too few retrieval results, and some documents with slightly lower relevance but still valuable are excluded; if the threshold is set too low, the retrieval results may include many irrelevant documents, affecting the accuracy of the retrieval. The calculated cosine similarity is compared with the preset value, and the documents corresponding to the document vectors with a cosine similarity greater than the preset value are screened out. These documents are considered to have a high semantic relevance to the query information and meet the user's retrieval requirements. For example, if the preset value is 0.7, then when the cosine similarity between a certain document vector and the query vector is greater than 0.7, the document corresponding to the document vector will be added to the vector retrieval results. Forming the vector retrieval results: After screening, all eligible documents constitute the vector retrieval results. These documents will serve as the basis for subsequent processing, be fused with the keyword retrieval results, and provide a rich information source for generating the final answer.

[0108] Optionally, in some embodiments of the present application, the step of "performing keyword retrieval on the updated query information" may specifically include:

[0109] Extract keywords from the updated query information;

[0110] Calculate the term weights and term frequencies of the extracted keywords in each document of the target document library;

[0111] Based on the term weights and terms, as well as the document lengths of each document in the target document library, calculate the relevance between each document in the target document library and the updated query information;

[0112] Based on the relevance, determine the documents associated with the updated query information in the target document library.

[0113] For example, specifically, through lexical analysis, the query information is segmented into individual words, stop words such as "de", "le", "zai" which have no substantial retrieval value are removed, and then key content words such as nouns and verbs are selected as keywords. Taking "How enterprises should address data security challenges in digital transformation" as an example, keywords such as "enterprise", "digital transformation", "data security challenges", "address" can be extracted. These keywords condense the core content of the query information and are the key basis for subsequent retrieval. Then, the term weights and term frequencies of the extracted keywords in each document of the target document library are calculated. Among them, the term weights are often calculated using the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm. The term frequency (TF) refers to the number of times a keyword appears in a certain document. The more frequently it appears, the higher the TF value. The inverse document frequency (IDF) reflects the importance of the keyword and is obtained by calculating the reciprocal of the proportion of the number of documents containing the keyword in the total number of documents and taking the logarithm. If a keyword appears in fewer documents, its IDF value is high, indicating that the word has a high discrimination degree and greater retrieval value.

[0114] For example, in a target document library containing numerous enterprise management documents, "digital transformation" appears frequently in some documents (high TF value), and the proportion of documents related to this topic in the total document library is small (high IDF value), so its TF-IDF value is high. Then, the TF-IDF value (i.e., the term weight) and term frequency of each keyword will be calculated for each document in the target document library, providing a data basis for subsequent calculation of document relevance.

[0115] In addition, the document length will affect the actual meaning of the keywords. To eliminate this influence, the system normalizes the term weights and term frequencies. Specifically, the term frequency can be divided by the document length to obtain the normalized term frequency, and then considering the term weight and normalized term frequency of the keyword comprehensively, the relevance score between the document and the query information is calculated through algorithms such as weighted summation. Suppose document B contains keywords "enterprise" and "data security challenges", and the term weights are w3 and w4 respectively, and the normalized term frequencies are tf3 and tf4 respectively. Then the relevance score SB of document B and the query information is SB = w3 × tf3 + w4 × tf4. In this way, the degree of relevance between the document and the query information can be evaluated more accurately. Based on the relevance, documents associated with the updated query information are determined in the target document library: After calculating the relevance score, the system sorts the documents according to the score. Usually, a threshold is set, and only documents with scores higher than the threshold will be recognized as associated with the query information. For example, if the threshold is set to 0.6, documents with a relevance score greater than 0.6 will be selected as the keyword retrieval results.

[0116] Optionally, in some embodiments of the present application, it may further include:

[0117] Perform chunking processing on the documents in the target document library;

[0118] Generate the embedding vector of each document block through the pre-trained language model;

[0119] The generated document blocks and embedding vectors are transmitted to the edge device through the network, so that after the edge device receives the data, it stores the generated document blocks and embedding vectors in the local cache according to the set cache strategy.

[0120] In some embodiments of the present application, the documents in the target document library can be preprocessed so that the edge devices can efficiently store and use these data. The documents in the target document library are often long and complex in content, and the direct processing efficiency is low. Block processing is to divide a long document into smaller document blocks, each of which contains relatively independent information. For example, a long document on the review of artificial intelligence technology may be divided into document blocks introducing different topics such as the history of artificial intelligence development, core algorithms, and application fields. The purpose of block division is to reduce the complexity of document processing and improve retrieval efficiency. During retrieval, the system can locate the specific content blocks related to the query more quickly without having to search the entire long document, which also facilitates the subsequent targeted analysis and processing of the document content.

[0121] Each document block is converted into an embedding vector through a pre-trained language model. This vector is a digital representation of the semantic information of the document block. Each dimension in the vector contains information about the document block on different semantic features. For example, a document block about "medical imaging diagnosis technology" will generate an embedding vector that reflects the characteristics of semantics related to medical care, imaging, diagnosis, etc. in certain dimensions after being processed by a pre-trained language model. In this way, the document block in text form is converted into a vector form that can be efficiently processed by computers, laying the foundation for subsequent vector-based retrieval and matching operations.

[0122] In order to reduce the delay in edge devices obtaining data in real time and improve the system response speed, it is necessary to transfer the processed document blocks and embedded vectors to the edge devices and cache them. After the data is transmitted to the edge devices through the network, the edge devices store it according to the pre-set cache strategy. The cache strategy can be formulated based on factors such as the frequency of use and importance of the data. For example, the document blocks and vectors in frequently queried fields are cached first, or the least recently used (LRU) algorithm is used to remove the least recently accessed document blocks and vectors from the cache to make room for new data. In this way, when a user initiates a query, the edge device can quickly obtain relevant data directly from the local cache for retrieval and processing, which greatly improves the response efficiency of the system, while also reducing the dependence on network connections, and ensuring a certain quality of service even when the network is unstable.

[0123] 105. Based on the preset prompt words and the processed document, generate a target answer corresponding to the query information through the preset large language model, and display the target answer through a digital human.

[0124] Before starting to generate an answer, it is necessary to clarify the preset prompt words and prepare the processed document first. The preset prompt words are the key instructions to guide the large language model to generate answers that meet the requirements, and they will be customized according to different fields, question types or user needs. In the Q&A in the legal field, the prompt words may be "Analyze the responsibilities of all parties in the following case according to relevant laws and regulations and give legal basis"; in the medical field, it may be "Combine medical knowledge to diagnose this symptom and give treatment suggestions". The processed document is the content highly relevant to the query information obtained by screening and sorting from the mixed retrieval results, and contains the key knowledge points and information required to answer the question.

[0125] In actual processing, input the preset prompt words, the processed document and the original query information into the preset large language model together. The large language model will understand and analyze these input information. Among them, the prompt words provide the direction and framework for the model to answer, the processed document provides specific knowledge content, and the original query information clarifies the core of the question. For example, when the user queries "What matters need to be noted when buying a second-hand house", the processed document contains relevant content such as the second-hand housing transaction process, property rights issues, and contract terms, and the preset prompt words are "According to the provided materials, concisely list the key points to note when buying a second-hand house". The large language model combines this information and starts to generate the answer content.

[0126] The large language model uses its own language understanding and generation capabilities to reason and create based on the input information. It will extract key information from the processed document and organize the language according to the requirements of the prompt words to generate a complete target answer. During the generation process, the model will consider the logic, coherence and accuracy of the language to make the answer meet the user's expectations as much as possible. For the above question about buying a second-hand house, the model may generate an answer like "Matters needing attention when buying a second-hand house: Confirm whether the property rights of the house are clear, check the real estate certificate and relevant certificates; Understand the actual situation of the house, such as the quality of the house, whether there is a mortgage, etc.; Carefully review the terms of the purchase contract to clarify the rights and obligations of both parties; Verify whether the transaction price of the house is reasonable, etc."

[0127] The generated target answer will be transmitted to the digital human interaction module. The digital human can display the answer content in two ways: voice and text. If it is a voice display, the digital human will use text-to-speech technology to convert the text answer into natural and fluent speech and convey it to the user in a vivid way; if it is a text display, the digital human will present the answer in text form on the interaction interface for the user to read conveniently. In the intelligent customer service scenario, the digital human will select the voice or text method to display the answer according to the user's preference, enhancing the user's interaction experience and enabling the user to obtain the required information more intuitively.

[0128] Optionally, in some embodiments of the present application, the step of "generating a target answer corresponding to the query information based on the preset prompt words and the processed document and displaying the target answer through the digital human" may specifically include:

[0129] Combining the query information, the preset prompt words, and the processed document to obtain an input set;

[0130] Inputting the input set into the preset large language model to generate a target answer that meets the prompt words;

[0131] Generating an interaction special effect for the target answer and displaying the target answer through the digital human according to the interaction special effect.

[0132] Among them, the query information is the question raised by the user, which clarifies the demand direction. The preset prompt words are the key instructions to guide the large language model to generate specific types of answers, such as "Please explain this concept in an easy-to-understand way" and "List three relevant practical cases", etc. It provides a framework and style guidance for the model's answer. The processed document is obtained after operations such as deduplication, screening, and sorting of the retrieved documents, and contains specific knowledge and content related to the query information. For example, when the user queries "What is blockchain technology", the preset prompt word is "Explain blockchain with simple life examples", and the processed document contains professional content such as the principle and characteristics of blockchain. Combining these three forms an input set, providing comprehensive input information for the large language model.

[0133] Based on its pre-trained parameters and algorithms, the large language model understands and analyzes the information in the input set. It will identify the intention of the query information, extract relevant knowledge and information from the processed document according to the requirements of the prompt words, and perform integration and reasoning.

[0134] During the answering process, the large language model will utilize the language knowledge and logical reasoning ability it has learned to organize language and ensure the accuracy, completeness, and coherence of the answer. For a query like "What is blockchain technology?", the large language model might generate an answer such as "Blockchain is like a ledger jointly maintained by everyone. For example, in a neighborhood, everyone records each transaction together, and each person has a complete ledger, making it difficult for anyone to tamper with the transaction records. Blockchain technology uses a similar method to make data more secure and transparent."

[0135] To make the process of the digital human presenting the answer more vivid and appealing, interactive effects for the target answer will be generated. The interactive effects can include the digital human's facial expression changes, gesture movements, adjustments in speech intonation, etc. When the digital human answers some important content, the key points can be highlighted through the emphasis of facial expressions or the assistance of gestures; when explaining complex concepts, speech effects such as appropriately slowing down the speech rate and intensifying the tone can be used to help users understand.

[0136] The digital human will perform corresponding voice output and action demonstrations according to the content of the answer and the requirements of the interactive effects. It will convey the target answer to the user in a natural and fluent manner, providing a good interactive experience and enabling users to obtain information more intuitively.

[0137] To further understand the information retrieval solution of this application, please refer to Figure 3 and Figure 4 , as Figure 3 shown, the embodiment of this application provides an artificial intelligence knowledge retrieval generation system based on digital human real-time interaction and edge device deployment (hereinafter referred to as the system).

[0138] In this embodiment, the system mainly consists of a digital human interaction module, a local device deployment module, a database management module, and a retrieval enhancement generation module. The digital human interaction module interacts with users via voice or text and obtains the user's query information in real time. After combining the query information with the domain information, it is transmitted to the local large language model. The local large language model generates a hypothetical answer based on the user's input and sends the answer along with the original query to the knowledge retrieval generation system for processing.

[0139] The edge device deploys a lightweight large language model and retrieves and infers the documents through a combination of hybrid retrieval such as vector retrieval and BM25 retrieval and RAG chain technology. Through the retrieval algorithm running locally on the edge device, relevant information can be quickly retrieved from the pre-loaded document library, and accurate answers can be generated using the retrieved relevant content, prompt words, and large language model. By combining with the digital human interaction module, the system not only reduces the dependence on the cloud, lowers network latency, but also improves the user experience.

[0140] As shown Figure 4 below: First, the user submits a query question through the digital human interaction interface. Then, based on the query information, its relevant field is judged. If there is no database in this field, a database cache is obtained by preprocessing the dataset in this field; if there is a database in this field, the database is called from the cache. Then, it is processed by a localized large language model according to the user's query to generate a hypothetical answer, and the hypothetical answer is combined with the original query to form a new query and sent to the retrieval module. In the retrieval module, two retrieval modes, vector retrieval and keyword retrieval, are performed on the new query to form a hybrid retrieval. Then, the content of the hybrid retrieval is deduplicated and sorted. Next, the sorted relevant documents are sent to the large language model, and the large language model is processed in combination with customized prompt words to generate the final answer. Finally, the generated answer is fed back to the user through the digital human interaction interface. By this way of combining digital human interaction and edge device deployment, the whole system not only improves the accuracy and relevance of knowledge retrieval, but also significantly improves the response speed and user experience.

[0141] Thus, it effectively solves problems such as the limitations of large language models, data privacy protection, and edge device computing resource limitations in professional field knowledge retrieval, fully exerts the advantages of edge device deployment, and improves the overall performance and efficiency of the intelligent knowledge retrieval system.

[0142] The above completes the information retrieval process of this application.

[0143] As can be seen from the above, this application provides an information retrieval method. After receiving the query information input by the user and determining the query database corresponding to the query information, a reference answer corresponding to the query information is generated based on the query database and the local model. Then, the query information and the reference answer are spliced to obtain the updated query information. Then, a hybrid retrieval is performed on the updated query information, and the documents obtained from the hybrid retrieval are processed. Finally, based on the preset prompt words and the processed documents, the target answer corresponding to the query information is generated through the preset large language model, and the target answer is displayed through the digital human. In the information retrieval solution provided by this application, a reference answer corresponding to the query information can be generated based on the query database and the local model. On this basis, the query information and the reference answer are spliced, and a hybrid retrieval is performed on the spliced updated query information, so that the query intention can be understood more accurately, thereby improving the relevance and accuracy of the retrieval results.

[0144] To facilitate better implementation of the information retrieval method of this application, this application also provides an information retrieval device based on the above. The meaning of the nouns is the same as that in the above information retrieval method, and the specific implementation details can refer to the description in the method embodiment.

[0145] Please refer toFigure 5 , Figure 5 is a schematic structural diagram of the information retrieval device provided by this application. The information retrieval device may include a receiving module 201, a generating module 202, a splicing module 203, a retrieval module 204, and a display module 205, which are specifically as follows:

[0146] The receiving module 201 is configured to receive the query information input by the user and determine the query database corresponding to the query information;

[0147] The generating module 202 is configured to generate a reference answer corresponding to the query information based on the query database and the local model;

[0148] The splicing module 203 is configured to splice the query information and the reference answer to obtain the updated query information;

[0149] The retrieval module 204 is configured to perform hybrid retrieval on the updated query information and process the documents obtained by the hybrid retrieval;

[0150] The display module 205 is configured to generate a target answer corresponding to the query information through a preset large language model based on the preset prompt words and the processed documents, and display the target answer through a digital human.

[0151] Optionally, in some embodiments of this application, the receiving module 201 may specifically be configured to: identify the keywords in the query information; determine the query database corresponding to the query information based on the string matching strategy and the keywords.

[0152] Optionally, in some embodiments of this application, the receiving module 201 may specifically further be configured to: if no keywords in the query information are identified, generate a query vector corresponding to the query information; obtain the average vector corresponding to each preset database; determine the query database corresponding to the query information according to the similarity between the query vector and each average vector.

[0153] Optionally, in some embodiments of this application, the retrieval module 204 may specifically be configured to: extract the query vector corresponding to the updated query information; perform retrieval on the updated query information based on the query vector, and perform keyword retrieval on the updated query information; fuse the vector retrieval result and the keyword retrieval result.

[0154] Optionally, in some embodiments of this application, the retrieval module 204 may specifically be configured to: input the query vector into a preset vector database, and calculate the cosine similarity between the query vector and each document vector in the vector database; add the documents corresponding to the document vectors with the cosine similarity greater than the preset value to the vector retrieval result.

[0155] Optionally, in some embodiments of the present application, the retrieval module 204 may specifically be configured to: extract keywords from the updated query information; calculate the term weights and term frequencies of the extracted keywords in each document of the target document library; calculate the relevance between each document of the target document library and the updated query information based on the term weights and terms, as well as the document lengths of each document in the target document library; and determine the documents associated with the updated query information in the target document library based on the relevance.

[0156] Optionally, in some embodiments of the present application, the display module 205 may specifically be configured to: combine the query information, the preset prompt words, and the processed documents to obtain an input set; input the input set into a preset large language model to generate a target answer that meets the prompt words; generate an interactive effect for the target answer, and display the target answer through a digital human according to the interactive effect.

[0157] The above completes the information retrieval process of the present application.

[0158] As can be seen from the above, the present application provides an information retrieval device. After the receiving module 201 receives the query information input by the user and determines the query database corresponding to the query information, the generating module 202 generates a reference answer corresponding to the query information based on the query database and the local model. Then, the splicing module 203 splices the query information and the reference answer to obtain the updated query information. Then, the retrieval module 204 performs hybrid retrieval on the updated query information and processes the documents obtained from the hybrid retrieval. Finally, the display module 205 generates a target answer corresponding to the query information through a preset large language model based on the preset prompt words and the processed documents, and displays the target answer through a digital human. In the information retrieval solution provided by the present application, a reference answer corresponding to the query information can be generated based on the query database and the local model. On this basis, the query information and the reference answer are spliced, and hybrid retrieval is performed on the spliced updated query information, so as to more accurately understand the query intention, thereby improving the relevance and accuracy of the retrieval results.

[0159] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructions or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0160] The embodiment of the present invention also provides an electronic device 500, such as Figure 6As shown, the electronic device 500 may integrate the above information retrieval device, and may further include a radio frequency (RF) circuit 501, a memory 502 including one or more computer-readable storage media, an input unit 503, a display unit 504, sensors 505, an audio circuit 506, a wireless fidelity (WiFi) module 507, a processor 508 including one or more processing cores, and a power supply 509 and other components. Those skilled in the art can understand that Figure 6 the structure of the electronic device 500 shown in Figure 6 does not limit the electronic device 500, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0161] The RF circuit 501 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is handed over to one or more processors 508 for processing; in addition, the data related to the uplink is sent to the base station. Generally, the RF circuit 501 includes but is not limited to antennas, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 501 can also communicate with the network and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to the global system of mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short messaging service (SMS), etc.

[0162] The memory 502 can be used to store software programs and modules. The processor 508 executes various functional applications and information processing by running the software programs and modules stored in the memory 502. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a target data playback function, etc.); the data storage area can store data created according to the use of the electronic device 500 (such as audio data, phone book, etc.). In addition, the memory 502 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 502 can also include a memory controller to provide access to the memory 502 by the processor 508 and the input unit 503.

[0163] The input unit 503 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, in a specific embodiment, the input unit 503 can include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch-sensitive surface), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 508, and can receive and execute commands sent by the processor 508. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit 503 can also include other input devices. Specifically, the other input devices can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), a trackball, a mouse, a joystick, etc.

[0164] The display unit 504 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device 500. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 504 may include a display panel. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, a touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor 508 to determine the type of touch event. Subsequently, the processor 508 provides a corresponding visual output on the display panel according to the type of touch event. Although in Figure 4 the touch-sensitive surface and the display panel are implemented as two independent components to achieve input and input functions, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.

[0165] The electronic device 500 may further include at least one sensor 505, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or the backlight when the electronic device 500 is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. As for other sensors that the electronic device 500 can also be configured with, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., they will not be elaborated here.

[0166] The audio circuit 506, the speaker, and the microphone can provide an audio interface between the user and the electronic device 500. The audio circuit 506 can convert the received audio data into an electrical signal and transmit it to the speaker, which converts it into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 506, converted into audio data, and then the audio data is output to the processor 508 for processing and then transmitted through the RF circuit 501 to, for example, another electronic device 500, or the audio data is output to the memory 502 for further processing. The audio circuit 506 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.

[0167] WiFi belongs to short-range wireless transmission technology. Through the WiFi module 507, the electronic device 500 can help users send and receive emails, browse the web, and access streaming media, etc. It provides users with wireless broadband Internet access. Although Figure 4 the WiFi module 507 is shown, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted entirely within the scope of not changing the essence of the invention according to needs.

[0168] The processor 508 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 502, and by calling the data stored in the memory 502, it executes various functions of the electronic device 500 and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 508 may include one or more processing cores; preferably, the processor 508 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 508 either.

[0169] The electronic device 500 also includes a power supply 509 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 508 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 509 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power data indicator.

[0170] Although not shown, the electronic device 500 may also include a camera, a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 508 in the electronic device 500 will, according to the following instructions, load the executable files corresponding to the processes of one or more application programs into the memory 502, and the processor 508 will run the application programs stored in the memory 502 to realize various functions:

[0171] Receive the query information input by the user and determine the query database corresponding to the query information; generate a reference answer corresponding to the query information based on the query database and the local model; splice the query information and the reference answer to obtain the updated query information; perform a hybrid retrieval on the updated query information and process the documents obtained from the hybrid retrieval; generate a target answer corresponding to the query information based on the preset prompt words and the processed documents through a preset large language model, and display the target answer through a digital human.

[0172] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For the parts not described in detail in a certain embodiment, reference may be made to the detailed description of the above information retrieval method, which will not be elaborated here.

[0173] As can be seen from the above, the electronic device 500 according to the embodiment of the present invention can detect whether the basic corpus data has a corresponding subtitle file. When the basic corpus data does not have a subtitle file, it performs speech alignment on the basic corpus data according to its built-in subtitle information to obtain the target corpus data, without the need for manual cleaning and annotation of the corpus data, which can not only reduce the cost of corpus collection, but also improve the efficiency of corpus collection.

[0174] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0175] Therefore, the embodiment of the present application further provides a storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute the steps in the above information retrieval method.

[0176] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0177] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0178] Since the instructions stored in the storage medium can execute the steps in any one of the information retrieval methods provided by the embodiments of the present invention, the beneficial effects that can be achieved by any one of the information retrieval methods provided by the embodiments of the present invention can be realized. For details, reference may be made to the previous embodiments, which will not be elaborated here.

[0179] The above has introduced in detail the information retrieval method, device, system and storage medium provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An information retrieval method, characterized in that, including: Receiving the query information input by the user and determining the query database corresponding to the query information; Generating a reference answer corresponding to the query information based on the query database and the local model; Concatenating the query information and the reference answer to obtain the updated query information; Performing hybrid retrieval on the updated query information and processing the documents obtained from the hybrid retrieval; Generating a target answer corresponding to the query information based on a preset prompt and the processed documents through a preset large language model, and presenting the target answer through a digital human.

2. The information retrieval method according to claim 1, wherein The performing hybrid retrieval on the updated query information and processing the documents obtained from the hybrid retrieval includes: Extracting the query vector corresponding to the updated query information; Performing retrieval on the updated query information based on the query vector, and; Performing keyword retrieval on the updated query information; Fusing the vector retrieval result and the keyword retrieval result.

3. The information retrieval method according to claim 2, wherein The performing retrieval on the updated query information based on the query vector includes: Inputting the query vector into a preset vector database and calculating the cosine similarity between the query vector and each document vector in the vector database; Adding the documents corresponding to the document vectors with the cosine similarity greater than the preset value to the vector retrieval result.

4. The information retrieval method according to claim 2, wherein The performing keyword retrieval on the updated query information includes: Extracting keywords from the updated query information; Calculating the term weights and term frequencies of the extracted keywords in each document of the target document library; Calculating the relevance between each document in the target document library and the updated query information based on the term weights and terms, and the document lengths of each document in the target document library; Determining the documents associated with the updated query information in the target document library based on the relevance.

5. The information retrieval method according to claim 4, wherein It also includes: Performing chunking processing on the documents in the target document library; Generating embedding vectors for each document chunk through a pre-trained language model; Transmitting the generated document chunks and embedding vectors to the edge device through the network, so that after receiving the data, the edge device stores the generated document chunks and embedding vectors in the local cache according to the set caching policy.

6. The information retrieval method according to any one of claims 1 to 5, characterized in that The generating a target answer corresponding to the query information based on a preset prompt and the processed documents through a preset large language model, and presenting the target answer through a digital human includes: Combining the query information, the preset prompt, and the processed documents to obtain an input set; Inputting the input set into a preset large language model to generate a target answer that meets the prompt; Generating interactive special effects for the target answer and presenting the target answer through a digital human according to the interactive special effects.

7. The information retrieval method according to any one of claims 1 to 5, characterized in that The determining the query database corresponding to the query information includes: Identifying the keywords in the query information; Determining the query database corresponding to the query information based on the string matching strategy and the keywords.

8. The information retrieval method according to claim 7, wherein It also includes: If no keywords in the query information are identified, generating the query vector corresponding to the query information; Obtaining the average vector corresponding to each preset database; Determine the query database corresponding to the query information according to the similarity between the query vector and each average vector.

9. An information retrieval device, characterized in that, Including: A receiving module, configured to receive the query information input by the user and determine the query database corresponding to the query information; A generating module, configured to generate a reference answer corresponding to the query information based on the query database and the local model; A splicing module, configured to splice the query information and the reference answer to obtain the updated query information; A retrieval module, configured to perform hybrid retrieval on the updated query information and process the documents obtained by the hybrid retrieval; A display module, configured to generate a target answer corresponding to the query information through a preset large language model based on the preset prompt words and the processed documents, and display the target answer through a digital human.

10. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the information retrieval method according to any one of claims 1 to 8 are implemented.

11. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, the steps of the information retrieval method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Scheme retrieval method and device based on industry large model, equipment and medium

    CN120744098A

  • Processing method and device for converting natural language into structured query language and medium

    CN122240655A