Large language model-based information search system, smart glasses and information search method
By introducing a large language model information search system into smart glasses, the problem of limited functionality in smart glasses has been solved, enabling more efficient information retrieval, improving search accuracy and real-time performance, and enhancing the intelligence level of smart glasses.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOLOS TECH SHENZHEN LTD
- Filing Date
- 2025-10-18
- Publication Date
- 2026-05-21
AI Technical Summary
Existing smart glasses have limited functionality, lack search capabilities, and are insufficient in terms of accuracy, real-time performance, and convenience of information retrieval.
An information search system based on a large language model is adopted. Through the collaborative work of smart glasses and a model server, speech-to-text and text-to-speech conversion is achieved. The large language model is used to determine the characteristics of user questions and obtain accurate answers.
It improves the real-time performance, completeness, and accuracy of information retrieval for smart glasses, and enhances the intelligence and interactivity of smart glasses.
Smart Images

Figure CN2025128597_21052026_PF_FP_ABST
Abstract
Description
Information search system, smart glasses, and information search methods based on large language models
[0001] This application claims priority to Chinese Patent Application No. CN 2024116344262, filed on November 15, 2024, entitled "Information Search System, Smart Glasses and Information Search Method Based on Large Language Model", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of smart glasses technology, and in particular to an information search system based on a large language model, smart glasses, and an information search method. Background Technology
[0003] With the development of computer technology, smart glasses are becoming more and more popular. However, existing smart glasses are expensive, and apart from their basic functions as smart glasses, they usually only have the functions of listening to music and making or receiving phone calls. Their functions are relatively simple and their level of intelligence is low.
[0004] Adding search functionality to smart glasses would make them more popular with users. However, improving the accuracy, completeness, and real-time nature of search results is a question worthy of in-depth research. Technical issues
[0005] The embodiments of this application aim to provide an information search system, smart glasses, and information search method based on a large language model, which can improve the real-time performance, completeness, convenience, and accuracy of information search based on the smart glasses system, and enhance the intelligence and interactivity of the smart glasses system. Technical solutions
[0006] One embodiment of this application provides an information search system based on a large language model, including a smart glasses system and a model server, wherein a large language model is configured on the model server;
[0007] The smart glasses system is used to acquire a first voice containing a user's question to be queried, acquire prompt information, convert the first voice into first text, and send the prompt information and the first text to a model server.
[0008] The model server is used to determine the characteristics of the user's question based on the prompt information and the first text using the large language model, obtain a second text containing the answer to the question based on the characteristics, and send the second text to the smart glasses system.
[0009] The smart glasses system is also used to convert the second text into second speech and play it.
[0010] One aspect of this application also provides a smart glasses based on a large language model. The smart glasses include: a frame, temples, at least one microphone, at least one speaker, a processor, and a memory, wherein the temples are connected to the frame, and the processor is connected to the at least one microphone, the at least one speaker, and the memory.
[0011] The memory stores a computer program that can be executed by the processor. The computer program includes multiple instructions that, when executed by the processor, cause the processor to:
[0012] A first voice recording containing the user's question to be queried is acquired through the at least one microphone;
[0013] The first speech is converted into the first text using a speech-to-text engine;
[0014] The characteristics of the user's question are determined by the large language model based on the prompt information and the first text, and a second text containing the answer to the question is obtained based on the characteristics.
[0015] The second text is converted into second speech using a text-to-speech engine, and the second speech is played through the at least one speaker.
[0016] This application also provides an information search method based on a large language model, applied to a smart mobile terminal, the information search method including:
[0017] The system receives a first voice message containing a user's question to be queried, sent by a smart wearable device, and converts the first voice message into first text using a speech-to-text engine, wherein the speech-to-text engine is configured on the smart mobile terminal or a speech-to-text server.
[0018] The characteristics of the user's question are determined based on the prompt information and the first text using a large language model, and a second text containing the answer to the question is obtained based on the characteristics, wherein the large language model is configured on the smart mobile terminal or model server.
[0019] The second text is converted into second speech using a text-to-speech engine, and the second speech is sent to the smart wearable device for playback, wherein the text-to-speech engine is configured in the smart mobile terminal or a text-to-speech server. Beneficial effects
[0020] The embodiments of this application utilize a large language model to realize information search based on natural language speech in smart glasses (or smart wearable devices), thereby improving the real-time performance, completeness, convenience, and accuracy of information search based on smart glasses (or smart wearable devices) or their collaboration with smart mobile terminals. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of smart glasses (or smart wearable devices) can be further improved. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a schematic diagram of the structure of an information search system based on a large language model provided in an embodiment of this application;
[0023] Figure 2 is an architectural schematic of an information search system based on a large language model provided in an embodiment of this application;
[0024] Figure 3 shows another architectural intent of an information search system based on a large language model provided in an embodiment of this application;
[0025] Figure 4 shows another architectural intent of an information search system based on a large language model provided in an embodiment of this application;
[0026] Figure 5 shows another architectural intent of an information search system based on a large language model provided in an embodiment of this application;
[0027] Figure 6 is a schematic diagram of the external structure of a smart glasses based on a large language model according to an embodiment of this application;
[0028] Figure 7 is a schematic diagram of the internal structure of the smart glasses shown in Figure 6;
[0029] Figure 8 is a flowchart illustrating the implementation of an information search method based on a large language model according to an embodiment of this application. Embodiments of the present invention
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0032] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0033] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0034] Referring to Figure 1, Figure 1 is a schematic diagram of the structure of an information search system based on a large language model provided in an embodiment of this application. As shown in Figure 1, the information search system 100 includes: a smart glasses system 110 and a model server 120, wherein a large language model (LLM) is configured on the model server 120.
[0035] The smart glasses system 110 is used to acquire a first voice containing a user's question to be queried, acquire a prompt message, convert the first voice into first text, and send the prompt message and the first text to the model server 120.
[0036] Model server 120 is used to determine the characteristics of the user's question based on the prompt information and the first text through the large language model, obtain a second text containing the answer to the question based on the characteristics, and send the second text to smart glasses system 110;
[0037] The smart glasses system 110 is also used to convert the second text into a second speech and play it.
[0038] The characteristics of a user question refer to whether the response is immediate, specifically including immediate (or real-time) and non-immediate (or non-real-time) responses. User questions with immediate characteristics have responses that are not fixed, but rather periodic, irregular, or change rapidly; examples include questions related to stock prices, news, daily weather, and the current time. User questions with non-immediate characteristics have responses that are fixed, unchanging over a long period, or not easily changed; examples include questions related to humanities, geography, and history.
[0039] The aforementioned large language models include: Generative Artificial Intelligence Large Language Model (GAILLM) and / or Multimodal Large Language Model (MLLM).
[0040] This generative AI large language model can be, for example, but is not limited to: OpenAI's ChatGPT, Google's Bard, and other models with similar functionality. This multimodal large language model can be, for example, but is not limited to: BLIP-2, LLaVA, MiniGPT-4, mPLUG-Owl, LLaMA-Adapter-v2, Otter, Multimodal-GPT, InstructBLIP, VisualGLM-6B, PandaGPT, LaVIN, and other models with similar functionality.
[0041] This prompt message is used to guide the large language model on how to obtain the answer to the user's question.
[0042] Optionally, in other embodiments of this application, the model server 120 is further configured to: extract multiple keywords for indicating user intent from the first text using the large language model, and obtain at least one sub-question corresponding to the multiple keywords;
[0043] Based on the prompt, the multiple keywords, and the at least one sub-question, determine the characteristic, and based on the characteristic, determine whether a search using a tool is necessary;
[0044] If not required, the system searches a pre-defined knowledge database for first response information matching the at least one sub-question, and generates the second text based on the first response information; and
[0045] If necessary, based on the prompt information, the multiple keywords, and the at least one sub-question, determine at least one corresponding search tool and search parameters, use the search parameters to call the at least one search tool to perform a search operation, obtain second response information matching the at least one sub-question based on the searched information, and generate the second text based on the second response information.
[0046] Specifically, answers to questions can be obtained by searching a pre-defined knowledge database or by using third-party tools to search from third-party online platforms. The appropriate method for obtaining the answer can be determined based on the characteristics of the user's question. Furthermore, the model server 120 can be configured with various third-party tool applications, such as, but not limited to, various search engines, web crawlers, vector database retrievers, weather apps, and clocks. The pre-defined knowledge database stores various types of knowledge information, such as, but not limited to, information related to astronomy, geography, history, humanities, and religion. This knowledge database can be configured on the model server 120 or on other cloud servers.
[0047] When the feature is non-real-time, it means that the answer to the question can be found in the preset knowledge database. Therefore, based on the at least one sub-question, the preset knowledge database is searched to obtain the first answer information that matches the at least one sub-question, and the second text is generated based on the first answer information.
[0048] When the feature is immediacy, it means that the answer to the question needs to be obtained by searching from a third-party network platform through a third-party tool. Therefore, based on the prompt information, the multiple keywords and the at least one sub-question, at least one search tool and search parameters required to obtain the answer to the question are determined. The at least one search tool is called by using the search parameters to perform the search operation. The second answer information matching the at least one sub-question is obtained based on the searched information, and the second text is generated based on the second answer information.
[0049] Optionally, in other embodiments of this application, the model server 120 is further configured to convert the searched information into text information through the large language model, cut the text information into text blocks and perform embedding calculations to generate text block vectors, and store the generated vectors in a vector database.
[0050] The model server is also used to retrieve the second response information from the vector database using the large language model.
[0051] Specifically, the large language model can use Retrieval-augmented Generation (RAG) to convert the searched information into text information, then cut the text information into text blocks and perform embedding calculations to generate text block vectors, and store the generated vectors in a vector database, and then retrieve the second response information from the vector database.
[0052] Optionally, in other embodiments of this application, the model server 120 is further configured to generate user session records based on the first text and the second text and store them in a historical session database;
[0053] The model server 120 is also used to search for historical session records associated with the user from the historical session database when it is determined that a search by a tool is needed, and to obtain the context information of the user's question based on the searched historical session records.
[0054] When the context information is related to at least one of the multiple keywords, the second response information is obtained by searching the vector database;
[0055] When the context information is not related to the multiple keywords, the corresponding at least one search tool and the search parameters are determined based on the prompt information, the multiple keywords and the at least one sub-question. The search operation is performed by calling the at least one search tool using the search parameters, and the second response information is obtained based on the search results.
[0056] Optionally, in other embodiments of this application, after the second voice is played, the smart glasses system 110 is further configured to acquire a third voice containing user instructions, convert the third voice into third text, and send the third text to the model server 120.
[0057] The model server 120 is also used to execute the task indicated by the user instruction based on the answer to the question through the large language model, generate a fourth text containing the execution result information of the task, and send it to the smart glasses system 110.
[0058] The smart glasses system 110 is also used to convert the fourth text into a fourth speech and play it.
[0059] Optionally, in other embodiments of this application, the information search system 100 may further include a data processing server 130 (not shown in FIG1).
[0060] The smart glasses system 110 is also used to obtain the prompt information through the data processing server 130, convert the first speech into the first text, and convert the second text into the second speech;
[0061] The smart glasses system 110 is also used to convert the third speech into the third text and the fourth text into the fourth speech via the data processing server 130.
[0062] Optionally, in other embodiments of this application, the smart glasses system 110 may include: smart glasses 111.
[0063] Optionally, in other embodiments of this application, the smart glasses system 110 may include: smart glasses 111 and smart mobile terminal 112;
[0064] Among them, the smart glasses 111 are used to acquire the first voice and send it to the smart mobile terminal 112, play the second voice, acquire the third voice and send it to the smart mobile terminal 112, and play the fourth voice.
[0065] The intelligent mobile terminal 112 is used to acquire the prompt information, convert the first speech into the first text, and send the prompt information and the first text to the model server 120, and convert the third speech into the third text and send it to the model server 120;
[0066] The model server 120 is also used to send the second text to the smart mobile terminal 112, and to send the fourth text to the smart mobile terminal 112;
[0067] The smart mobile terminal 112 is also used to convert the second text into the second speech and send it to the smart glasses 111, and to convert the fourth text into the fourth speech and send it to the smart glasses 111.
[0068] The structure and function of the smart glasses 111 can be specifically described in the embodiments shown in Figures 6 and 7 below, and will not be repeated here.
[0069] The intelligent mobile terminal 112 may include, but is not limited to, cellular phones, smartphones, other wireless communication devices, personal digital assistants, audio players, other media players, music recorders, video recorders, cameras, other media recorders, smart radios, laptops, personal digital assistants (PDAs), portable multimedia players (PMPs), Moving Image Experts Group (MPEG-1 or MPEG-2) Audio Layer 3 (MP3) players, digital cameras, and other intelligent devices capable of processing data while in motion. The intelligent mobile terminal may also have Android, iOS, or other operating systems installed.
[0070] Optionally, in other embodiments of this application, the smart glasses 111 or the smart mobile terminal 112 is further configured to generate prompt information through a prompt information generator, wherein the prompt information generator is configured in the smart glasses 111, the smart mobile terminal 112, or the data processing server 130.
[0071] Optionally, in other embodiments of this application, the prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines whether a search tool is needed, as well as the corresponding at least one search tool and the search parameters, based on the multiple keywords and the at least one sub-question.
[0072] These multiple application examples are application cases or examples in different application scenarios. By learning from these multiple application examples, the large language model can know which operations or tasks need to be performed based on the text sent by the smart glasses system 110, what information needs to be obtained, and what information needs to be fed back to the smart glasses system 110.
[0073] Optionally, in other embodiments of this application, the smart glasses 111 or the smart mobile terminal 112 is further configured to respond to a user's configuration command, obtain the configuration information indicated by the configuration command, and generate the prompt information according to the configuration information through the prompt information generator. The configuration information includes configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the retrieval conditions of the vector database.
[0074] The model server 120 can be configured with multiple different alternative large language models. The model server 120 is also used to determine the large language model from the different alternative large language models according to the prompt information, input the prompt information and the first text into the large language model, and obtain the second text output by the large language model. The large language model also configures its own parameters according to the configuration information.
[0075] Specifically, this application offers extensive configuration flexibility, allowing users to configure the large language model used in the interaction (i.e., users can select at least one from multiple preset large models such as ChatGPT 3.5 turbo, ChatGPT 4 turbo, Claude 3, Germini, etc.), information sources (e.g., Google, Wikipedia, or designated websites), available third-party tools (e.g., search engines, web crawlers, vector database retrievers, weather forecasts, clocks, etc.), data processing programs (e.g., splitting, embedding, storage), and vector database retrieval criteria (e.g., maximum number of documents, minimum cosine similarity, or using multiple queries) to meet different user needs. This flexibility ensures that this application can adapt to application requirements of varying complexity and achieve fine-tuning for optimized performance.
[0076] Optionally, the data processing server 130 can be a single server or a server cluster consisting of multiple servers. When the data processing server 130 is a server cluster consisting of multiple servers, in this embodiment, the data processing server 130 may specifically include: a speech-to-text server 131 configured with a speech-to-text engine, a text-to-speech server 132 configured with a text-to-speech engine, a prompt server 133 configured with the prompt information generator, a knowledge storage server 134 configured with the knowledge database, a vector server 135 configured with the vector database, and a log server 136 configured with the historical session database. (Not shown in Figure 1)
[0077] The smart glasses system 110 is also used to send the first speech to the speech-to-text server 131, so that the speech-to-text server 131 converts the first speech into the first text through the speech-to-text engine and sends it to the smart glasses system 110.
[0078] The smart glasses system 110 is also used to send the second text to the text-to-speech server 132, so that the text-to-speech server 132 converts the second text into the second speech through the speech-to-text engine and sends it to the smart glasses system 110.
[0079] The smart glasses system 110 is also used to send a notification message to the vector server 135 after the user finishes the query;
[0080] The smart glasses system 110 is also used to obtain the prompt information through the prompt server 133;
[0081] Vector server 135 is also used to clear all vector data associated with the user in the vector database based on the notification information.
[0082] Specifically, the aforementioned information search system can have two different system architectures. Referring to Figures 2 and 3, as shown in Figure 2, in the first system architecture, the information search system 100 includes: smart glasses 111 and a cloud server 220, wherein the smart glasses 111 and the cloud server 220 interact via long-range wireless networks such as WiFi / 3 / 4 / 5G. Alternatively, as shown in Figure 3, in the second system architecture, the information search system 100 includes: smart glasses 111, a smart mobile terminal 112, and a cloud server 330, wherein the smart glasses 111 and the smart mobile terminal 112 interact via short-range low-power wireless networks such as Bluetooth, and the smart mobile terminal 112 and the cloud server 330 interact via long-range wireless networks such as WiFi / 3 / 4 / 5G.
[0083] The cloud server 220 or 330 can be a single server integrating a large language model, a prompt message generator, the aforementioned engines, and the aforementioned databases, or it can be a server cluster consisting of multiple servers. As shown in Figure 2, the cloud server 220 may include a data processing server 130 and a model server 120.
[0084] As an extension, as shown in Figure 4, the information search system can also have a third system architecture. In this third system architecture, the information search system 100 includes: smart glasses 111, a smart mobile terminal 112, and a data processing server 130. The data processing server 130 includes: a model server 120, a speech-to-text server 131, a text-to-speech server 132, a prompt server 133, a knowledge storage server 134, a vector server 135, and a log server 136. The smart glasses 111 and the smart mobile terminal 112 interact via a short-range low-power wireless network such as Bluetooth. The smart glasses 111 or the smart mobile terminal 112 interacts with the data processing server 130 via a long-range wireless network such as WiFi / 3G / 4G / 5G.
[0085] The following will provide a detailed explanation of the information search process of the information search system 100, using examples 1-3 applied to three different scenarios.
[0086] Example 1 (Weather Inquiry)
[0087] In this application example 1, assuming the user's question is "What's the weather like in Redmond, Oregon today?", the information search system 100 will perform the following operations:
[0088] 1. The smart glasses system 110 picks up the query voice from the microphone and uses a speech-to-text (STT) engine to convert the query voice into text;
[0089] 2. The smart glasses system 110 sends the text and corresponding prompts from the STT engine to the large language model LLM configured on the model server 120 to obtain the answer to the user's question;
[0090] 3. Based on the received prompts and text, LLM performs the following operations to obtain an answer to the user's question;
[0091] 3-1 Understanding Queries
[0092] (1) Obtain the topic (keyword 1): weather;
[0093] (2) Get specific details (keyword 2): location (Redmond, Oregon) and date (today).
[0094] 3-2. Decompose the query
[0095] Obtaining subproblems:
[0096] (1) What is the weather like in Redmond, Oregon right now?
[0097] (2) What is the weather forecast for the rest of the day in Redmond, Oregon?
[0098] 3-3. Determine whether a tool is needed (LLM will collect information related to the characteristics of the user's question, such as the keywords "today," "weather," and "remaining time" mentioned above, and make a decision on whether a tool is needed based on the collected information. In this application example 1, the question about today's weather is immediate, so a tool is needed.)
[0099] (1) Select the search tool: "Weather Forecast" tool;
[0100] (2) Specify search parameters: location (Redmond, Oregon), date (today).
[0101] 3-4. Use the search tool to perform the search operation.
[0102] Use the specified search parameter "Redmond, Oregon, today" to invoke the Weather Forecast tool to obtain weather data (from the results of the Weather Forecast tool).
[0103] 3-5. General Information
[0104] Filter and extract useful information from the weather data returned by the "Weather Forecast" tool, such as current weather data and weather forecast data.
[0105] (1) Current weather data and weather forecast data are combined into a detailed answer to the question;
[0106] (2) Including temperature, humidity, wind speed and major weather events.
[0107] 4. The LLM sends a text containing the answer to the question to the smart glasses system 110, in which the answer clearly and in detail presents the weather information;
[0108] 5. The smart glasses system 110 converts the detailed weather information (i.e., the answer to the question) in the received text into speech through a text-to-speech (TTS) engine, and then outputs it to the speaker of the smart glasses 111 for playback.
[0109] A specific example of the answer to the above question could be: "The weather in Redmond, Oregon today is as follows: the current temperature is 25°C, the humidity is 60%, the wind speed is 10 kilometers per hour, and it is expected to be sunny for most of the day, with a small chance of rain in the afternoon."
[0110] Example 2 (Stock Price Inquiry)
[0111] In this application example 2, assuming the user's question is "What is the current stock price of NVIDIA?", the information search system 100 will perform the following operations:
[0112] 1. The smart glasses system 110 picks up the voice used to query the above user questions through the microphone of the smart glasses, and uses the STT engine to convert the voice into text;
[0113] 2. The smart glasses system 110 sends the text and corresponding prompts from the STT engine to the large language model LLM configured on the model server 120 to obtain the answer to the user's question;
[0114] 3. Based on the received prompts and text, LLM performs the following operations to obtain an answer to the user's question;
[0115] 3-1 Understanding Queries
[0116] (1) Obtain the main topic (keyword 1): current stock price;
[0117] (2) Obtain specific details (keyword 2): Company (NVIDIA, NVDA).
[0118] 3-2. Decompose the query
[0119] Sub-question: What is Nvidia's current stock price?
[0120] 3-3. Determine whether a tool is needed (LLM will collect information related to the characteristics of the user's question, such as the keywords "stock price" and "current" mentioned above, and make a decision on whether a tool is needed based on the collected information. In this application example 2, the question about the current stock price is immediate, so a tool is needed.)
[0121] (1) Select a search tool: Search engine tool
[0122] (2) Specify search parameters: query("Nvidia's current stock price"), global relevance
[0123] 3-4. Use the search tool to perform the search operation.
[0124] Use specified search parameters (e.g., "query('Nvidia's current stock price'), global relevance") to invoke a search engine tool (e.g., the Google API) to obtain the latest stock price data.
[0125] 3-5. Result Embedding
[0126] The search results are embedded into vectors and stored in a vector database, and documents are split, embedded into vectors, and then stored in that vector database.
[0127] Specifically, the search results obtained by the search engine tool are converted into text information, the text information is cut into text blocks and embedded to generate text block vectors, and then the generated vectors are stored in a vector database.
[0128] Preferably, since the user's question is immediate, in order to save space and improve the utilization of storage space, all data related to this conversation in the vector database will be released after the conversation ends.
[0129] 3-6. Data Retrieval
[0130] (1) Extract relevant stock price information from the vector database and generate a question answer based on the extracted information, wherein only documents relevant to the context of the user's question are retrieved;
[0131] (2) Ensure that the data is up-to-date and accurate.
[0132] 4. The LLM sends a text containing the answer to the question to the smart glasses system 110, in which the answer clearly and in detail presents information about Nvidia's stock price;
[0133] Specifically, this may include, but is not limited to, the date and time of the last update of Nvidia's stock price and any significant market movements.
[0134] 5. The smart glasses system 110 converts the stock price information (from LLM) in the received text into speech through a TTS engine, and then outputs it to the speaker of the smart glasses 111 for playback.
[0135] A specific response to the above question could be: "According to the latest data, as of 1:00 p.m. on July 29, 2024, Nvidia's (NVDA) current stock price is $450.00. Please note that stock prices are subject to change and may vary throughout the trading day."
[0136] Furthermore, if a user raises a new request after listening to the voice response to the question, the smart glasses system 110 can also use LLM to perform corresponding operations based on the new request.
[0137] For example, after playing a voice message containing searched Nvidia stock price information, if the system detects that the user has issued a voice message containing a buy / sell instruction for Nvidia stock, the smart glasses system 110 uses STT (Simplified Translation Time) to convert the voice message into text, and then sends the converted text to the LLM (Limited Module). Based on the previously searched Nvidia stock price information, the LLM calls third-party stock trading software to execute the buy / sell Nvidia stock task, generates a fourth text message containing the execution result information of the task, and sends it to the smart glasses system 110. The smart glasses system 110 then uses TTS (Text-to-Speech) to convert this fourth text message into a fourth voice message and plays it. This execution result information may include, but is not limited to, notifications of whether the buy / sell of Nvidia stock was successful, the Nvidia stock buy / sell price, and the buy / sell time.
[0138] In another application example of this application, the user can also use the smart glasses system 110 to query flights for a certain day. After playing the query results, if the user issues a voice command to book a ticket, the smart glasses system 110 can also use LLM to call a dedicated third-party ticketing tool to perform the ticketing task.
[0139] Example 3 (Historical Query)
[0140] In this application example 3, assuming the user's question is "Why did the Roman Empire decline?", the information search system 100 will perform the following operations:
[0141] 1. The smart glasses system 110 picks up the query voice from the microphone and uses the STT engine to convert the query voice into text;
[0142] 2. The smart glasses system 110 sends the text and corresponding prompts from the STT engine to the large language model LLM configured on the model server 120 to obtain the answer to the user's question;
[0143] 3. Based on the received prompts and text, LLM performs the following operations to obtain an answer to the user's question;
[0144] 3-1 Understanding Queries
[0145] Topic (Keyword 1): The Decline of the Roman Empire
[0146] Obtain detailed information (keyword 2): Historical reasons for the fall
[0147] 3-2. Decompose the query
[0148] Obtaining subproblems:
[0149] (1) What were the main reasons for the fall of the Roman Empire?
[0150] (2) When did the Roman Empire fall?
[0151] 3-3. Determine whether tools are needed (LLM will collect information related to the characteristics of the user's question, such as the keywords "Rome," "destruction," "decline," "history," and "cause" mentioned above, and make a decision on whether tools are needed based on the collected information. In this application example 3, the question about the reasons for the decline of the Roman Empire is not immediate, so tools are not needed.)
[0152] We determined to utilize existing historical event knowledge bases to obtain answers to the questions.
[0153] 3-4. General Information
[0154] By searching a historical event knowledge base, we obtained the key reasons and timeline for the decline of the Roman Empire, and then combined these key reasons and timelines into a coherent question answer.
[0155] 4. The LLM sends a text containing the answers to the questions to the smart glasses system 110, in which the answers clearly and in detail present the timeline of the decline of the Roman Empire and the various factors that led to its decline.
[0156] A specific answer to the above question could be: "The decline of the Roman Empire was a complex historical event with multiple contributing factors. The Western Roman Empire officially collapsed in 476 AD, when its last emperor, Romulus Augustulus, was deposed by the Germanic chieftain OdoArthur. The key factors leading to its decline were mainly as follows: 1. Economic decline: Heavy taxes, dependence on slave labor, and economic recession weakened the empire; 2. Military problems: The Roman legions constantly faced threats from invading barbarian tribes and struggled to maintain their vast borders; 3. Political corruption: Poor leadership, political instability, and corruption weakened the government's ability to manage the empire; 4. Social problems: Declining civic pride and public services, as well as widening wealth inequality, led to social unrest; 5. External pressure: Continuous invasions by various barbarian groups, including the Visigoths, Vandals, and Huns, placed enormous pressure on the empire's defenses. These factors ultimately led to the final collapse of the Western Roman Empire, marking the end of ancient Rome and the beginning of the Middle Ages."
[0157] 5. The smart glasses system 110 converts the received text-based question answers (from LLM) into speech using a TTS engine, and then outputs them to the speaker of the smart glasses 111 for playback.
[0158] For ease of understanding, the following will use the architecture shown in Figure 5 as an example to further explain the above-mentioned information search system based on a large language model. As shown in Figure 5, the audio input from the user picked up by the smart glasses 111 is transmitted to the STT engine for conversion into text. Then, the text and prompt information are sent by the smart glasses 111 to the data processing server 130.
[0159] The data processing server 130 can be a single server or a distributed server cluster consisting of multiple servers, including the model server 120, speech-to-text server 131, text-to-speech server 132, prompt server 133, knowledge storage server 134, vector server 135, and log server 136. Understandably, in practical applications, depending on actual needs, the data processing server 130 may include some or all of the aforementioned servers, or it may include more servers with other functions.
[0160] The data processing server 130 can provide users with personalized information search services. On the server side, it assigns a user an associated account with a customized virtual assistant. Based on the user's custom operations, the server can configure one or more of the following search conditions for different users' virtual assistants: large language models, tools, information sources, knowledge databases, data processing methods, and vector databases. This is to meet the different search needs of different users, thereby making the search more targeted, flexible, and accurate.
[0161] Specifically, this virtual assistant allows each LLM to interact with external APIs (Application Programming Interfaces), such as LLM Wrapper, to access user-customized large language models (LLMs). These LLMs then identify the user's intent based on the text and prompts, determining whether to retrieve a question answer from a pre-defined knowledge database. LLM Wrapper can be used to access numerous large language models, such as OpenAI, Cohere, and Hugging Face.
[0162] If the LLM outputs a result that the answer to the question can be retrieved from the knowledge database, the virtual assistant will retrieve the answer from the knowledge database, thus shortening the search time and improving the response speed.
[0163] Otherwise, if an online search is determined to be necessary, the virtual assistant runs tools provided by the information search service, such as a search engine, to retrieve real-time data. The retrieved real-time data is broken down and embedded into vectors, which are then stored in a vector database. The LLM retrieves relevant information from this vector database and returns it to the virtual assistant. This process iterates until the virtual assistant accumulates enough information to respond to the user. The LLM's output may include instructions for invoking multiple tools, allowing the virtual assistant to run multiple tools simultaneously to gather more information in parallel. The virtual assistant's response is then streamed to the user to enhance the user experience.
[0164] Ultimately, the text-based answers were converted to speech using TTS and read aloud through the smart glasses. Furthermore, the chat log was saved to provide the virtual assistant with contextual awareness throughout the conversation.
[0165] In the embodiments of this application, a large language model is used in the smart glasses system to realize real-time information search based on natural language speech, thereby improving the real-time performance, convenience and accuracy of information search based on the smart glasses system. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses system can be further improved.
[0166] Referring to Figures 6 and 7, Figure 6 is a schematic diagram of the external structure of a smart glasses based on a large language model according to an embodiment of this application, and Figure 7 is a schematic diagram of the internal structure of the smart glasses shown in Figure 6. As shown in Figures 6 and 7, the smart glasses 111 includes: a frame 1111, at least one temple 1112 (two temples are shown in Figure 6 for ease of understanding), at least one microphone 1113, at least one speaker 1114, a processor 1115, and a memory 1116, wherein the temple 1112 is connected to the frame 1111, and the processor 1115 is connected to at least one microphone 1113, at least one speaker 1114, and the memory 1116. It is understood that Figures 6 and 7 are only examples, and in practical applications, the smart glasses 111 may have more or fewer components or parts than shown in Figures 6 and 7 as needed.
[0167] The memory 1116 stores a computer program that can be executed by the processor 1115. This computer program includes multiple instructions that, when executed by the processor 1115, cause the processor 1115 to:
[0168] Acquire a first voice containing the user's question to be queried through at least one microphone 1113;
[0169] The first speech is converted into the first text using a speech-to-text engine.
[0170] The characteristics of the user's question are determined by the prompt information and the first text using a large language model, and a second text containing the answer to the question is obtained based on these characteristics.
[0171] The second text is converted into second speech using a text-to-speech engine, and the second speech is played through at least one speaker 1114.
[0172] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 further causes to: extract a plurality of keywords for indicating user intent from the first text through the large language model, and obtain at least one sub-question corresponding to the plurality of keywords; obtain the prompt information through the prompt information generator, and determine the feature based on the prompt information, the plurality of keywords and the at least one sub-question, and determine whether it is necessary to search using a tool based on the feature; if not, search for first answer information matching the at least one sub-question from a preset knowledge database, and generate the second text based on the first answer information; and if necessary, determine at least one corresponding search tool and search parameters based on the prompt information, the plurality of keywords and the at least one sub-question, call the at least one search tool to perform a search operation using the search parameters, obtain second answer information matching the at least one sub-question based on the searched information, and generate the second text based on the second answer information.
[0173] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to acquire, after playing the second speech through at least one speaker 1114, a third speech containing user instructions through at least one microphone 1113, and convert the third speech into third text.
[0174] Based on the answer to the question, the task indicated by the user instruction is executed using this large language model, and a fourth text containing the execution result information of the task is generated.
[0175] The fourth text is converted into a fourth speech, and the fourth speech is played through at least one speaker 1114.
[0176] In other embodiments of this application, the large language model is configured on smart glasses 111, smart mobile terminal 112, or model server 120; the speech-to-text engine is configured on smart glasses 111, smart mobile terminal 112, or speech-to-text server 131; the text-to-speech engine is configured on smart glasses 111, smart mobile terminal 112, or text-to-speech server 132; and the prompt information generator is configured on smart glasses 111, smart mobile terminal 112, or prompt server 133.
[0177] In other embodiments of this application, the smart glasses 111 further includes a Bluetooth communication component 1117. When the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to send the first voice to the smart mobile terminal 112 via the Bluetooth communication component 1117, and instruct the smart mobile terminal 112 to: convert the first voice into the first text via the speech-to-text engine; obtain the second text based on the prompt information and the first text via the large language model configured on the smart mobile terminal 112 or the model server 120; and convert the second text into the second voice via the text-to-speech engine, wherein the prompt information is generated by the smart mobile terminal 112 via a prompt information generator configured on the smart mobile terminal 112 or the prompt server 133; and receive the second voice sent by the smart mobile terminal 112 via the Bluetooth communication component 1117.
[0178] Furthermore, when the processor 1115 executes the multiple instructions, it also causes the processor 1115 to send the third voice to the smart mobile terminal 112 via the Bluetooth communication component 1117, and instruct the smart mobile terminal 112 to convert the third voice into the third text via the speech-to-text engine, execute the task indicated by the user instruction based on the question answer using the large language model configured in the smart mobile terminal 112 or the model server 120, and generate a fourth text containing the execution result information of the task, and convert the fourth text into the fourth voice via the text-to-speech engine; and receive the fourth voice sent by the smart mobile terminal 112 via the Bluetooth communication component 1117.
[0179] Optionally, in other embodiments of this application, the smart glasses 111 further includes a wireless communication component 1118. When the plurality of instructions are executed by the processor 1115, the processor 1115 also causes the processor 1115 to send the first speech to the speech-to-text server 131 via the wireless communication component 1118, so that the first speech is converted into the first text by the speech-to-text engine in the speech-to-text server; to obtain the prompt information generated by the prompt information generator on the prompt server 133 from the prompt server 133 via the wireless communication component 1118; to send the first text and the prompt information to the model server 120, so that the second text is obtained by the large language model in the model server 120 based on the prompt information and the first text; and to send the second text to the text-to-speech server 132, so that the second text is converted into the second speech by the text-to-speech engine in the text-to-speech server 132.
[0180] Furthermore, when the processor 1115 executes the multiple instructions, it also causes the processor 1115 to send the third speech to the speech-to-text server 131 via the wireless communication component 1118, so that the speech-to-text engine in the speech-to-text server 131 can convert the third speech into the third text; send the third text to the model server 120, so that the large language model in the model server 120 can execute the task indicated by the user instruction based on the question answer and generate a fourth text containing the execution result information of the task; send the fourth text to the text-to-speech server 132, so that the text-to-speech engine in the text-to-speech server 132 can convert the fourth text into the fourth speech.
[0181] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to convert the searched information into text information through the large language model, cut the text information into text blocks and perform embedding calculations to generate text block vectors, store the generated vectors in a vector database, and retrieve the second response information from the vector database.
[0182] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 also causes the processor 1115 to clear all vector data associated with the user in the vector database after the user ends the query.
[0183] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to generate the user's session record based on the first text and the second text and store it in the historical session database; when the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to, through the large language model,: when it is determined that a search is required by a tool, search the historical session database for historical session records associated with the user, and obtain the context information of the user's question based on the searched historical session records; when the context information is related to at least one of the plurality of keywords, obtain the second response information by searching the vector database; and when the context information is not related to any of the plurality of keywords, determine the corresponding at least one search tool and the search parameters based on the prompt information, the plurality of keywords and the at least one sub-question, call the at least one search tool to perform the search operation by using the search parameters, and obtain the second response information based on the searched information.
[0184] Optionally, in other embodiments of this application, the prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines whether a search tool is needed, as well as the corresponding at least one search tool and the search parameters, based on the multiple keywords and the at least one sub-question.
[0185] Optionally, in other embodiments of this application, when the plurality of instructions are executed by the processor 1115, the processor 1115 is also caused to: respond to the user's configuration instruction, obtain the configuration information indicated by the configuration instruction; generate the prompt information according to the configuration information through the prompt information generator, the prompt information further including the configuration information, the configuration information including configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the retrieval conditions of the vector database.
[0186] Furthermore, the smart glasses 111 also includes a battery 210 (not shown in Figure 7) for providing power to the various electronic devices of the smart glasses 111, such as at least one microphone 1113, at least one speaker 1114, processor 1115, memory 1116, Bluetooth communication module 1117, and wireless communication module 1118.
[0187] The various electronic components of the aforementioned smart glasses can be connected via a bus.
[0188] It should be noted that the components of the smart glasses described above can be either substitutes or superimposed. That is, all the components described in this embodiment can be installed on a single smart pair of glasses, or, depending on the requirements, only a portion of the components can be installed. When the substitution relationship is present, the smart glasses also have a peripheral connection interface. This connection interface can be at least one of the following: PS / 2 interface, serial interface, parallel interface, IEEE1394 interface, USB (Universal Serial Bus) interface, etc. The function of the substituted component can be realized through a peripheral connected to this connection interface, such as an external speaker or external sensor.
[0189] For further details regarding the smart glasses in this embodiment, please refer to the relevant descriptions in the embodiments shown in Figures 1 to 5 and Figure 8, which will not be repeated here.
[0190] In the embodiments of this application, information search based on natural language speech is realized by using a large language model through smart glasses, thereby improving the real-time performance, completeness, convenience and accuracy of information search based on smart glasses. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of smart glasses can be further improved.
[0191] Referring to Figure 8, which is a flowchart illustrating the implementation of an information search method based on a large language model according to an embodiment of this application, applied to the smart mobile terminal 112 in the above embodiment, the information search method includes:
[0192] S801. Receive a first voice message containing a user question to be queried sent by a smart wearable device, and convert the first voice message into first text using a speech-to-text engine, wherein the speech-to-text engine is configured on the smart mobile terminal or a speech-to-text server.
[0193] S802. Using a large language model, determine the characteristics of the user's question based on the prompt information and the first text, and obtain a second text containing the answer to the question based on the characteristics, wherein the large language model is configured on the smart mobile terminal or model server.
[0194] S803. The second text is converted into second speech through a text-to-speech engine, and the second speech is sent to the smart wearable device for playback, wherein the text-to-speech engine is configured in the smart mobile terminal or a text-to-speech server.
[0195] In this embodiment, the smart wearable device may include, but is not limited to, smart helmets, smart headphones, smart earrings, smartwatches, smart glasses, and other wearable smart devices.
[0196] Optionally, in other embodiments of this application, in step S802, the characteristics of the user's question are determined based on the prompt information and the first text using a large language model, and a second text containing the question's answer is obtained based on these characteristics. Specifically, this includes:
[0197] S8021. Using the large language model, extract multiple keywords from the first text to indicate the user's intent, and obtain at least one sub-question corresponding to the multiple keywords;
[0198] S8022. Based on the prompt information, the multiple keywords, and the at least one sub-question, determine the characteristic, and based on the characteristic, determine whether it is necessary to search using a tool;
[0199] S8023. If not required, search a preset knowledge database for first response information matching the at least one sub-question, and generate the second text based on the first response information, wherein the knowledge database is configured in the smart mobile terminal or knowledge storage server; and
[0200] S8024. If necessary, determine at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords and the at least one sub-question, call the at least one search tool to perform a search operation using the search parameters, obtain second response information matching the at least one sub-question based on the searched information, and generate the second text based on the second response information.
[0201] Optionally, in other embodiments of this application, tasks related to the question answer can be executed according to subsequent user instructions, such as buying tickets, buying / selling stocks, sharing the question answer, generating related articles and publishing them on online social platforms, etc., thereby giving the smart wearable device more functions. Specifically, after sending the second voice to the smart wearable device for playback, the information search method further includes: receiving a third voice containing user instructions sent by the smart wearable device, and converting the third voice into third text using the speech-to-text engine; executing the task indicated by the user instruction based on the question answer using the large language model, and generating a fourth text containing the execution result information of the task; and converting the fourth text into fourth voice using the text-to-speech engine, and sending the fourth voice to the smart wearable device for playback.
[0202] Optionally, in other embodiments of this application, step S8024, obtaining second response information matching the at least one sub-question based on the searched information, includes: converting the searched information into text information, cutting the text information into text blocks and performing embedding calculations to generate text block vectors, and storing the generated vectors in a vector database, wherein the vector database is configured in the smart mobile terminal or vector server; and retrieving the second response information from the vector database.
[0203] Optionally, in other embodiments of this application, the information search method further includes: after the user ends the query, clearing all vector data associated with the user in the vector database.
[0204] Optionally, in other embodiments of this application, the information search method further includes: generating the user's session records based on the first text and the second text and storing them in a historical session database, wherein the historical session database is configured in the smart mobile terminal or log server.
[0205] Before determining at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords, and the at least one sub-question, the information search method further includes: searching the historical session database for historical session records associated with the user, obtaining contextual information of the user's question based on the searched historical session records; and obtaining the second response information by searching the vector database when the contextual information is related to at least one of the multiple keywords. The step of determining at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords, and the at least one sub-question includes: when the contextual information is not related to any of the multiple keywords, determining at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords, and the at least one sub-question.
[0206] Optionally, in other embodiments of this application, the prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines whether a search tool is needed, as well as the corresponding at least one search tool and the search parameters, based on the multiple keywords and the at least one sub-question.
[0207] Optionally, in other embodiments of this application, the information search method further includes: obtaining configuration information indicated by a user's configuration instruction; and sending the configuration information to a prompt information generator, so that the prompt information generator generates a prompt information containing the configuration information, wherein the configuration information includes configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the search conditions of the vector database. The prompt information generator can be configured on a smart mobile terminal or a prompt server.
[0208] Optionally, in other embodiments of this application, before obtaining the second text containing the answer to the question through the large language model based on the prompt information and the first text, the information search method further includes: obtaining the prompt information through the prompt information generator.
[0209] For further details regarding the information search method based on large language models, please refer to the relevant descriptions in the embodiments shown in Figures 1 to 7, which will not be repeated here.
[0210] In the embodiments of this application, by collaborating with a smart mobile terminal and a smart wearable device, and utilizing a large language model, information search based on natural language speech is realized. This improves the real-time performance, completeness, convenience, and accuracy of information search based on the smart mobile terminal and the smart wearable device. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart wearable device can be further enhanced.
[0211] This application also provides a non-transitory computer-readable storage medium, which may be disposed in the smart glasses or smart wearable devices described in the above embodiments. This non-transitory computer-readable storage medium may be the memory 304 in the embodiment shown in FIG. 7. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the information search method based on a large language model described in the above embodiments. Furthermore, the computer-readable storage medium may also be a USB flash drive, a portable hard drive, a read-only memory (ROM), RAM, a magnetic disk, or an optical disk, or any other medium capable of storing program code.
[0212] In the embodiments provided in this application, it should be understood that the disclosed information search system, smart glasses, and information search method based on a large language model can be implemented in other ways. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the interconnections shown or discussed, whether direct or indirect, can be through interfaces, devices, or modules, and can be electrical, mechanical, or other forms.
[0213] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0214] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0215] The above is a description of the information search system, smart glasses, and information search method based on a large language model provided in this application. For those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An information retrieval system based on a large language model, characterized in that, The system includes a smart glasses system and a model server, wherein a large language model is configured on the model server. The smart glasses system is used to acquire a first voice containing a user's question to be queried, acquire prompt information, convert the first voice into first text, and send the prompt information and the first text to a model server. The model server is used to determine the characteristics of the user's question based on the prompt information and the first text using the large language model, obtain a second text containing the answer to the question based on the characteristics, and send the second text to the smart glasses system. The smart glasses system is also used to convert the second text into second speech and play it.
2. The information search system as described in claim 1, characterized in that, The model server is also used for: Using the large language model, multiple keywords indicating user intent are extracted from the first text, and at least one sub-question corresponding to the multiple keywords is obtained; Based on the prompt information, the multiple keywords, and the at least one sub-question, determine the characteristic, and based on the characteristic, determine whether a search using a tool is necessary; If not required, the first response information matching the at least one sub-question is searched from a preset knowledge database, and the second text is generated based on the first response information; as well as If necessary, at least one search tool and search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question. The search parameters are used to call the at least one search tool to perform a search operation. The search results are used to obtain a second response information that matches the at least one sub-question. The second text is then generated based on the second response information.
3. The information search system as described in claim 2, characterized in that, The model server is also used to convert the searched information into text information through the large language model, cut the text information into text blocks and perform embedding calculations to generate text block vectors, and store the generated vectors in a vector database. The model server is also used to retrieve the second response information from the vector database using the large language model.
4. The information search system as described in claim 3, characterized in that, The model server is also used to generate user session records based on the first text and the second text and store them in the historical session database; The model server is also used to, through the large language model, When it is determined that a search using a tool is necessary, the historical session records associated with the user are searched from the historical session database, and the context information of the user's question is obtained based on the searched historical session records. When the context information is related to at least one of the plurality of keywords, the second response information is obtained by searching the vector database; When the context information is not related to the multiple keywords, the corresponding at least one search tool and the search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question. The search operation is performed by calling the at least one search tool using the search parameters, and the second response information is obtained based on the searched information.
5. The information search system as described in claim 4, characterized in that, After the second voice message finishes playing The smart glasses system is also used to acquire third voice containing user commands, convert the third voice into third text, and send the third text to the model server; The model server is also used to execute the task indicated by the user instruction based on the question answer using the large language model, generate a fourth text containing the execution result information of the task, and send it to the smart glasses system. The smart glasses system is also used to convert the fourth text into a fourth speech and play it.
6. The information search system as described in claim 5, characterized in that, The information search system also includes a data processing server; The smart glasses system is also used to obtain the prompt information through the data processing server, convert the first voice into the first text, and convert the second text into the second voice; The smart glasses system is also used to convert the third speech into the third text and the fourth text into the fourth speech via the data processing server.
7. The information search system as described in claim 6, characterized in that, The smart glasses system includes: smart glasses; or, The smart glasses system includes: smart glasses and a smart mobile terminal; wherein, the smart glasses are used for: The system acquires the first voice message and sends it to the smart mobile terminal, plays the second voice message, acquires the third voice message and sends it to the smart mobile terminal, and plays the fourth voice message. The intelligent mobile terminal is used to acquire the prompt information, convert the first voice into the first text, and send the prompt information and the first text to the model server, and convert the third voice into the third text and send it to the model server; The model server is also configured to send the second text to the smart mobile terminal, and to send the fourth text to the smart mobile terminal. The smart mobile terminal is also configured to convert the second text into the second speech and send it to the smart glasses, and to convert the fourth text into the fourth speech and send it to the smart glasses.
8. The information search system as described in claim 7, characterized in that, The smart glasses or the smart mobile terminal are further configured to generate the prompt information via a prompt information generator, wherein the prompt information generator is configured in the smart glasses, the smart mobile terminal, or the data processing server.
9. The information search system as described in claim 4, characterized in that, The prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines, based on the multiple keywords and the at least one sub-question, whether a search tool is needed and determines the corresponding at least one search tool and the search parameters.
10. The information search system as described in claim 8, characterized in that, The smart glasses or the smart mobile terminal are also used to respond to the user's configuration command, obtain the configuration information indicated by the configuration command, and generate the prompt information according to the configuration information through the prompt generator. The configuration information includes configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the retrieval conditions of the vector database. The model server is configured with multiple different candidate large language models. The model server is also used to determine the large language model from the different candidate large language models according to the prompt information, input the prompt information and the first text into the large language model, and obtain the second text output by the large language model. The large language model also configures its own parameters according to the configuration information.
11. The information search system as described in claim 10, characterized in that, The data processing server includes: a speech-to-text server configured with a speech-to-text engine, a text-to-speech server configured with a text-to-speech engine, a prompt server configured with the prompt information generator, a knowledge storage server configured with the knowledge database, a vector server configured with the vector database, and a log server configured with the historical session database. The smart glasses system is further configured to send the first speech to the speech-to-text server, so that the speech-to-text server converts the first speech into the first text through the speech-to-text engine and sends it to the smart glasses system; The smart glasses system is further configured to send the second text to the text-to-speech server, so that the text-to-speech server converts the second text into the second speech through the speech-to-text engine and sends it to the smart glasses system; The smart glasses system is also used to send a notification message to the vector server after the user finishes the query; The vector server is also configured to clear all vector data associated with the user in the vector database according to the notification information.
12. A smart glasses based on a large language model, characterized in that, The smart glasses include: a frame, temples, at least one microphone, at least one speaker, a processor, and a memory, wherein the temples are connected to the frame, and the processor is connected to the at least one microphone, the at least one speaker, and the memory; The memory stores a computer program that can be executed by the processor. The computer program includes multiple instructions that, when executed by the processor, cause the processor to: A first voice recording containing the user's question to be queried is acquired through the at least one microphone; The first speech is converted into the first text using a speech-to-text engine; The characteristics of the user's question are determined by the large language model based on the prompt information and the first text, and a second text containing the answer to the question is obtained based on the characteristics. The second text is converted into second speech using a text-to-speech engine, and the second speech is played through the at least one speaker.
13. The smart glasses as described in claim 12, characterized in that, When the processor executes the plurality of instructions, it also causes the processor to: Using the large language model, multiple keywords indicating user intent are extracted from the first text, and at least one sub-question corresponding to the multiple keywords is obtained; The prompt information is obtained through a prompt information generator, and the characteristics are determined based on the prompt information, the multiple keywords, and the at least one sub-question. Based on the characteristics, it is determined whether a search using a tool is necessary. If not required, the first response information matching the at least one sub-question is searched from a preset knowledge database, and the second text is generated based on the first response information; as well as If necessary, at least one search tool and search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question. The search parameters are used to call the at least one search tool to perform a search operation. The search results are used to obtain a second response information that matches the at least one sub-question. The second text is then generated based on the second response information.
14. The smart glasses as described in claim 13, characterized in that, When the processor executes the plurality of instructions, it also causes the processor, after playing the second voice through the at least one speaker, A third voice containing user instructions is acquired through the at least one microphone, and the third voice is converted into third text; Based on the question answer, the user-instructed task is executed using the large language model, and a fourth text containing the execution result information of the task is generated. The fourth text is converted into a fourth speech, and the fourth speech is played through the at least one speaker.
15. The smart glasses as described in claim 14, characterized in that, The large language model is configured on the smart glasses, smart mobile terminal, or model server; the speech-to-text engine is configured on the smart glasses, smart mobile terminal, or speech-to-text server; the text-to-speech engine is configured on the smart glasses, smart mobile terminal, or text-to-speech server; and the prompt information generator is configured on the smart glasses, smart mobile terminal, or prompt server. The smart glasses also include a Bluetooth communication component, which, when the processor executes the plurality of instructions, also enables the processor to communicate via the Bluetooth communication component. The first voice is sent to the smart mobile terminal, and the smart mobile terminal is instructed to: convert the first voice into the first text through the speech-to-text engine, obtain the second text based on the prompt information and the first text through the large language model configured on the smart mobile terminal or the model server, and convert the second text into the second voice through the text-to-speech engine, wherein the prompt information is obtained by the smart mobile terminal through the prompt information generator configured on the smart mobile terminal or the prompt server; as well as Receive the second voice message sent by the smart mobile terminal; as well as When the processor executes the plurality of instructions, it also causes the processor to communicate via the Bluetooth communication component. The third speech is sent to the smart mobile terminal, and the smart mobile terminal is instructed to: convert the third speech into the third text using the speech-to-text engine; execute the task indicated by the user instruction based on the question answer using the large language model configured on the smart mobile terminal or the model server, and generate a fourth text containing the execution result information of the task; and convert the fourth text into the fourth speech using the text-to-speech engine; and Receive the fourth voice message sent by the smart mobile terminal; or, The smart glasses also include a wireless communication component, which, when the processor executes the plurality of instructions, also causes the processor to communicate via the wireless communication component. The first speech is sent to the speech-to-text server so that the first speech is converted into the first text by the speech-to-text engine in the speech-to-text server; Obtain the prompt information generated by the prompt information generator on the prompt server; The first text and the prompt information are sent to the model server so that the second text can be obtained by the large language model in the model server based on the prompt information and the first text; as well as The second text is sent to the text-to-speech server so that the second text can be converted into the second speech by the text-to-speech engine in the text-to-speech server; and When the processor executes the plurality of instructions, it also causes the processor to communicate via the wireless communication component. The third speech is sent to the speech-to-text server so that the speech-to-text engine in the speech-to-text server can convert the third speech into the third text. The third text is sent to the model server so that the large language model in the model server can execute the task indicated by the user instruction based on the question answer and generate a fourth text containing the execution result information of the task. as well as The fourth text is sent to the text-to-speech server so that the text-to-speech engine in the text-to-speech server can convert the fourth text into the fourth speech.
16. The smart glasses as described in claim 15, characterized in that, When the processor executes the plurality of instructions, it also causes the processor to convert the searched information into text information through the large language model, cut the text information into text blocks and perform embedding calculations to generate text block vectors, store the generated vectors in a vector database, and retrieve the second response information from the vector database.
17. The smart glasses as described in claim 16, characterized in that, When the processor executes the plurality of instructions, it also causes the processor to clear all vector data associated with the user in the vector database after the user ends the query.
18. The smart glasses as described in claim 16, characterized in that, When the processor executes the plurality of instructions, it also causes the processor to generate the user's session record based on the first text and the second text and store it in the historical session database; When the processor executes the plurality of instructions, it also causes the processor to, through the large language model: When it is determined that a search using a tool is necessary, the historical session records associated with the user are searched from the historical session database, and the context information of the user's question is obtained based on the searched historical session records. When the context information is related to at least one of the plurality of keywords, the second response information is obtained by searching the vector database; as well as When the context information is not related to the multiple keywords, the corresponding at least one search tool and the search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question. The search operation is performed by calling the at least one search tool using the search parameters, and the second response information is obtained based on the searched information.
19. The smart glasses as described in claim 18, characterized in that, The prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines, based on the multiple keywords and the at least one sub-question, whether a search tool is needed and determines the corresponding at least one search tool and the search parameters.
20. The smart glasses as described in claim 19, characterized in that, When the processor executes the plurality of instructions, it also causes the processor to: In response to a user's configuration command, obtain the configuration information indicated by the configuration command; The prompt information generator generates the prompt information based on the configuration information. The prompt information also includes the configuration information, which includes configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the retrieval conditions of the vector database.
21. An information retrieval method based on a large language model, characterized in that, The information search method, applied to smart mobile terminals, includes: The system receives a first voice message containing a user's question to be queried, sent by a smart wearable device, and converts the first voice message into first text using a speech-to-text engine, wherein the speech-to-text engine is configured on the smart mobile terminal or a speech-to-text server. The characteristics of the user's question are determined based on the prompt information and the first text using a large language model, and a second text containing the answer to the question is obtained based on the characteristics, wherein the large language model is configured on the smart mobile terminal or model server. The second text is converted into second speech using a text-to-speech engine, and the second speech is sent to the smart wearable device for playback, wherein the text-to-speech engine is configured in the smart mobile terminal or a text-to-speech server.
22. The information search method as described in claim 21, characterized in that, The process of determining the characteristics of the user's question based on the prompt information and the first text using a large language model, and obtaining a second text containing the question's answer based on the characteristics, specifically includes: Using the large language model, multiple keywords indicating user intent are extracted from the first text, and at least one sub-question corresponding to the multiple keywords is obtained; Based on the prompt information, the multiple keywords, and the at least one sub-question, determine the characteristic, and based on the characteristic, determine whether a search using a tool is necessary; If not required, a first response matching the at least one sub-question is searched from a preset knowledge database, and the second text is generated based on the first response, wherein the knowledge database is configured in the smart mobile terminal or knowledge storage server; and If necessary, at least one search tool and search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question. The search parameters are used to call the at least one search tool to perform a search operation. The search results are used to obtain a second response information that matches the at least one sub-question. The second text is then generated based on the second response information.
23. The information search method as described in claim 22, characterized in that, After sending the second voice message to the smart wearable device for playback, the process further includes: Receive third voice containing user commands sent by a smart wearable device, and use the speech-to-text engine to convert the third voice into third text; Based on the question answer, the user-instructed task is executed using the large language model, and a fourth text containing the execution result information of the task is generated. The text-to-speech engine converts the fourth text into a fourth speech, and then sends the fourth speech to the smart wearable device for playback.
24. The information search method as described in claim 23, characterized in that, The step of obtaining second response information matching the at least one sub-question based on the searched information includes: The searched information is converted into text information, the text information is cut into text blocks and embedded to generate text block vectors, and the generated vectors are stored in a vector database, wherein the vector database is configured in the smart mobile terminal or vector server; The second response information is retrieved from the vector database.
25. The information search method as described in claim 24, characterized in that, The information search method also includes: After the user finishes the query, all vector data associated with the user is cleared from the vector database.
26. The information search method as described in claim 24, characterized in that, The information search method also includes: The user's session records are generated based on the first text and the second text and stored in a historical session database, wherein the historical session database is configured on the smart mobile terminal or log server; Before determining at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords, and the at least one sub-question, the information search method further includes: Search the historical session database for historical session records associated with the user, and obtain the context information of the user's question based on the searched historical session records; When the context information is related to at least one of the plurality of keywords, the second response information is obtained by searching the vector database; The step of determining at least one corresponding search tool and search parameters based on the prompt information, the multiple keywords, and the at least one sub-question includes: When the context information is not relevant to the multiple keywords, at least one corresponding search tool and search parameters are determined based on the prompt information, the multiple keywords, and the at least one sub-question.
27. The information search method as described in claim 26, characterized in that, The prompt information includes multiple application examples. The large language model learns from these multiple application examples and determines, based on the multiple keywords and the at least one sub-question, whether a search tool is needed and determines the corresponding at least one search tool and the search parameters.
28. The information search method as described in claim 27, characterized in that, The information search method also includes: In response to a user's configuration command, obtain the configuration information indicated by the configuration command; and The configuration information is sent to the prompt information generator so that the prompt information generator generates the prompt information containing the configuration information. The configuration information includes configuration parameters of at least one of the following: the large language model, the tool, the information source, the data processing method, and the retrieval conditions of the vector database. The prompt information generator is configured on the smart mobile terminal or the prompt server. Before obtaining the second text containing the answer to the question based on the prompt information and the first text using a large language model, the information search method further includes: The prompt information is obtained through the prompt information generator.