Intelligent glasses control system and method based on large language model and intelligent glasses
Through the smart glasses control system based on the large language model, the smart glasses work together with cloud servers and mobile terminals is realized, the problem of single function of smart glasses is solved, and the intelligence and interactivity of smart glasses are improved.
Patent Information
- Application Number
- CN202410034795.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing smart glasses have single functions, low intelligence and expensive.
The smart glasses control system based on the large language model is adopted, and the coordinated work of smart glasses, smart mobile terminals and cloud servers realizes the processing and interaction of images and voice, and uses the large language model to perform semantic analysis and reply generation.
It enriches the functions of smart glasses, improves intelligence and interactivity, and enhances the scalability and self-creation of smart glasses.
Smart Images

Figure CN120299454A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of smart glasses, and particularly to a control system for smart glasses based on a large language model, a smart glass, and a control method for smart wearable devices. Background Art
[0002] With the development of computer technology, smart glasses have become increasingly popular. However, existing smart glasses are expensive, and in addition to their inherent functions as smart glasses, they usually only have functions such as listening to music and making or answering phone calls, with relatively single functions and low intelligence levels. Summary of the Invention
[0003] Embodiments of the present application provide a control system for smart glasses based on a large language model, a smart glass, and a control method for smart wearable devices, which are used to enrich the functions of smart glasses and improve the intelligence and interactivity of smart glasses.
[0004] On the one hand, embodiments of the present application provide a control system for smart glasses based on a large language model, where the system includes: a smart glass, a smart mobile terminal, and a cloud server. The smart glass includes a microphone, a speaker, a camera, and a Bluetooth component, and the large language model is configured in the cloud server;
[0005] The smart glass is used to capture an image through the camera, obtain the user's first voice through the microphone, and send the image and the first voice to the smart mobile terminal through the Bluetooth component, where the first voice includes the question raised by the user;
[0006] The smart mobile terminal is used to send the first voice and the image to the cloud server;
[0007] The cloud server is used to: convert the first voice into a first text; perform semantic parsing on the first text and generate a prompt message according to the parsed semantics; through the large language model, obtain a second text according to the first text, the prompt message, and the image, where the second text contains the answer to the question; convert the second text into a second voice; and send the second voice to the smart mobile terminal;
[0008] The smart mobile terminal is further used to send the second voice to the smart glass;
[0009] The smart glass is further used to receive the second voice through the Bluetooth component and play the second voice through the speaker.
[0010] On the one hand, an embodiment of the present application further provides an intelligent glasses based on a large language model, including: a frame, at least one temple, a microphone, a speaker, at least one sensor, a processor, and a memory;
[0011] The at least one temple is connected to the frame, and the processor is electrically connected to the microphone, the speaker, the at least one sensor, and the memory;
[0012] One or more programs executable by the processor are stored in the memory. The one or more programs include a plurality of instructions, and the plurality of instructions are used for:
[0013] Obtain sensing data through the at least one sensor. Wherein, the at least one sensor includes a camera, and the sensing data includes an image captured by the camera;
[0014] Obtain the first voice of the user through the microphone, and the first voice includes the question raised by the user;
[0015] Through the large language model, according to the first voice and the sensing data, obtain a second voice including the answer to the question, wherein the large language model is configured in the intelligent glasses or an intelligent mobile terminal or a cloud server;
[0016] Play the second voice through the speaker.
[0017] On the one hand, an embodiment of the present application further provides a control method for an intelligent wearable device based on a large language model, which is applied to an intelligent mobile terminal. The method includes:
[0018] Receive the first voice and image sent by the intelligent wearable device through Bluetooth, wherein the first voice includes the question raised by the user;
[0019] Convert the first voice into a first text, perform semantic parsing on the first text, and generate a prompt message according to the parsed semantics;
[0020] Through the large language model, according to the image, the first text, and the prompt message, obtain a second text, wherein the second text includes the answer to the question, and the large language model is configured in the intelligent mobile terminal or a cloud server;
[0021] Convert the second text into a second voice, and send the second voice to the intelligent wearable device through the Bluetooth for playing.
[0022] In each embodiment of the present application, through the above control system using a large language model, human-machine task interaction based on captured images is realized on smart glasses or smart wearable devices, thus enriching the functions of the smart glasses. And due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses or smart wearable devices can be further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 FIG. is a schematic structural diagram of a smart glasses control system based on a large language model provided by an embodiment of the present application;
[0025] Figure 2 and Figure 3 is Figure 1 a schematic diagram of the application scenario of the shown control system;
[0026] Figure 4 FIG. is a schematic structural diagram of a smart glasses control system based on a large language model provided by another embodiment of the present application;
[0027] Figure 5 FIG. is a schematic internal structure diagram of smart glasses provided by an embodiment of the present application;
[0028] Figure 6 FIG. is a schematic external structure diagram of smart glasses provided by an embodiment of the present application;
[0029] Figure 7 FIG. is a flowchart of the implementation of a control method for a smart wearable device based on a large language model provided by an embodiment of the present application;
[0030] Figure 8 is Figure 7 a schematic diagram of Application Example 1 of the shown method;
[0031] Figure 9 is Figure 7 a schematic diagram of Application Example 2 of the shown method;
[0032] Figure 10 is Figure 7 a schematic diagram of Application Example 3 of the shown method;
[0033] Figure 11 is Figure 7 a schematic diagram of Application Example 4 of the shown method;
[0034] Figure 12 For Figure 7 a schematic diagram of Application Example 5 of the method shown Specific Embodiments
[0035] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0036] In the following, the terms "including", "having", and their cognates that can be used in various embodiments of the present invention are only intended to represent specific features, numbers, steps, operations, elements, components, or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing items.
[0037] In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0038] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which various embodiments of the present invention belong. The terms (such as those defined in a general-use dictionary) will be construed as having the same meaning as the contextual meaning in the relevant technical field and will not be construed as having an idealized meaning or an overly formal meaning unless clearly defined in various embodiments of the present invention.
[0039] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of a natural language command control system based on a generative artificial intelligence large language model provided by an embodiment of this application. As Figure 1 shown, the control system 100 includes: a smart glasses 110, a smart mobile terminal 120, and a cloud server 130.
[0040] Among them, the cloud server 130 may be a single server or a distributed server cluster composed of multiple servers.
[0041] The smart glasses 110 may be open smart glasses, including components such as a microphone, a speaker, a camera, and a Bluetooth component. For the specific structure of the smart glasses 110, reference can be made to the following Figure 5 andFigure 6 Related description of the illustrated embodiment.
[0042] The intelligent mobile terminal 120 may include, but is not limited to: a cellular phone, a smart phone, other wireless communication devices, a personal digital assistant, an audio player, other media players, a music recorder, a video recorder, a camera, other media recorders, a smart radio, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a moving picture expert group (MPEG-1 or MPEG-2) audio layer 3 (MP3) player, a digital camera, and an intelligent wearable device (such as a smart watch, a smart bracelet, etc.). The intelligent mobile terminal 120 is also installed with an Android, IOS or other operating system.
[0043] The smart glasses 110 are used to capture an image through a camera, obtain a first voice of a user through a microphone, and send the image and the first voice to the smart mobile terminal 120 through a Bluetooth component, wherein the first voice includes a question raised by the user, and the image can be a static picture, a dynamic video or image, or a combination of a picture and an image. The number of the images can be one or more.
[0044] The smart mobile terminal 120 is used to send the first voice and the image to the cloud server 130 .
[0045] The cloud server 130 is used to: convert the first speech into a first text; perform semantic analysis on the first text and generate prompt information based on the analyzed semantics; obtain a second text based on the first text, the prompt information and the image through a large language model (LLM), wherein the second text contains an answer to the question; convert the second text into a second speech; and send the second speech to the smart mobile terminal 120.
[0046] The smart mobile terminal 120 is further configured to send the second voice to the smart glasses 110 .
[0047] The smart glasses 110 are further configured to receive the second voice through the Bluetooth component and play the second voice through the speaker.
[0048] The large language model may be configured on the cloud server 130. The large language model may include, but is not limited to, a Generative Artificial Intelligence Large Language Model (GAILLM) or a Multimodal Large Language Model (MLLM).
[0049] The generative artificial intelligence large language model can be, for example, but not limited to: ChatGPT of Open AI, Bard of Google, and other models with similar functions. The multimodal large language model can be, for example, but not limited to: BLIP-2, LLaVA, MiniGPT-4, mPLUG-Owl, LLaMA-Adapter-v2, Otter, Multimodal-GPT, InstructBLIP, VisualGLM-6B, PandaGPT, LaVIN, and other models with similar functions. The large language model can include multiple parts, which are trained based on different samples to answer different task requests based on the image, which can be, but not limited to: image recognition, information sharing, making calls, navigation, translation, sending messages, sending emails, searching the network, calling network services, and calling other third-party software development tools (Software Development Kit, SDK) to perform tasks provided by the SDK, etc. Specific examples: user questions about "Where is this photo in Hong Kong?", "What am I looking at?", "Which direction should I go?", "Log in to the website", or "Can the text in this photo be translated?"
[0050] Taking translation as an example, when a user shoots a video (or one or more pictures) through the smart glasses 110 and speaks the first voice "Translate the English in the video into Chinese", the smart glasses 110 sends the video and the first voice to the smart mobile terminal 120 through the Bluetooth component, so that the video and the first voice are forwarded to the cloud server 130 through the smart mobile terminal 120.
[0051] After receiving the video and the first voice, the cloud server 130 converts the first voice into a first text, performs semantic analysis on the first text and generates prompt information according to the analyzed semantics, and then inputs the first text, the prompt information and the video into the large language model. The prompt information includes a translation task request.
[0052] The large language model extracts the text in the video (such as "The next station is Guangzhou South Railway Station") according to the first text and the prompt information, and translates the text using English as the source language and Chinese as the target language, and finally outputs a second text containing the translated content (such as "The next station is Guangzhou South Station"). Optionally, the source language can also be automatically identified by the large language model when performing a text extraction operation in the video.
[0053] The cloud server 130 converts the second text output by the large language model into second voice and forwards it to the smart glasses 110 through the smart mobile terminal 120 for playback.
[0054] Optionally, in another embodiment of the present application, the smart mobile terminal 120 is further configured to obtain location information, and send the location information, the first voice, and the image to the cloud server 130, where the location information is obtained by the smart mobile terminal 120 through a built-in satellite positioning system, or is obtained by the smart glasses 110 and sent to the smart mobile terminal 120. The cloud server 130 is further configured to, through the large language model, obtain the second text according to the location information, the first text, the prompt information, and the image. Wherein, the satellite positioning system may be GPS (Global Positioning System), or the satellite positioning system may also use Beidou satellites or other similar satellites for positioning.
[0055] Optionally, in another embodiment of the present application, the cloud server 130 is further configured to obtain historical chat records, and through the large language model, obtain the second text according to the historical chat records, the first text, the prompt information, and the image.
[0056] Optionally, in another embodiment of the present application, the smart mobile terminal 120 is further configured to obtain location information, and send the location information, the first voice, and the image to the cloud server 130, where the location information is obtained by the smart mobile terminal 120 through a built-in satellite positioning system, or is obtained by the smart glasses 110 and sent to the smart mobile terminal 120;
[0057] The cloud server 130 is further configured to obtain historical chat records, and through the large language model, obtain the second text according to the historical chat records, the location information, the first text, the prompt information, and the image.
[0058] As Figure 2 shown, in a practical application example, when a user is shopping and discovers a new building, the user can first take a photo through the camera on the worn smart glasses 110, and then say a question such as "What building is this" to the smart glasses 110. The smart glasses 110 obtain the voice of the user as the first voice through a built-in microphone, and then send the captured image and the obtained first voice to the smart phone 120 through Bluetooth.
[0059] The smart phone 120 obtains location information through a built-in positioning component, and then sends the location information, the received image, and the first voice to the cloud server 130. Wherein, the positioning component may be, but is not limited to, a positioning component based on GPS or Beidou satellites.
[0060] The cloud server 130 obtains the user's historical chat records, where the historical chat records include all texts and / or images exchanged among the smart glasses 110, the smartphone 120, and the cloud server 130 when the user chats with the large language model through the smart glasses 110 during the activation period of the chat function of the smart glasses 110 or within a preset period.
[0061] At the same time, the cloud server 130 converts the first voice into a first text, performs semantic parsing on the first text, and generates a prompt message based on the parsed semantics, where the prompt message contains the parsed semantics. Among them, as Figure 3 shown, in other application examples, according to the specific content of the parsed semantics, in addition to image recognition, the prompt message can also be used to prompt the large language model to perform at least one of the following tasks based on the image: information sharing, making a call, navigation, translation, sending a message, sending an email, searching the network, invoking a network service, and invoking other third-party SDKs to perform the tasks provided by the SDK. The control system uses the large language model through the cloud server 130 to process and interpret the user's voice commands in natural language, and then performs the tasks specified by the voice context like a personal assistant. These tasks can be to ask the model to find the name of a building after taking a photo with the camera, make a call, send a message, send an email, search for interests, invoke a network service to obtain / publish data, location, or navigation, and invoke other third-party SDKs to perform the tasks provided by the SDK. For example, after the user takes a photo of a building and makes a task request like "I want to know the name of this building", the large language model will answer the user's task request.
[0062] Then, the cloud server 130 inputs the historical chat records, the location information, the first text, the prompt message, and the image into the large language model, and obtains a second text output by the large language model. The second text contains the answer to the question. The cloud server 130 converts the second text into a second voice, and then sends the second voice to the smartphone 120, so that the smartphone 120 sends the second voice to the smart glasses 110 through Bluetooth for playback. Among them, the large language model uses the location information according to the prompt message to identify the building in the image, obtains the name of the building, and outputs a second text containing the name. Further, the second text may also include an introduction to the building, such as its origin, structure, allusions, etc.
[0063] In another application example, the user can first take a picture containing a QR code through the smart glasses 110 and say the first voice command "Log in to the website in the picture in the name of Jack with the password 123456". The smart glasses 110 send the picture and the first voice command to the cloud server 130 through the smart mobile terminal 120. The cloud server 130 converts the first voice command into the corresponding first text and performs semantic analysis on the first text to obtain a prompt message, where the prompt message includes: a task request to recognize the picture. Then, the cloud server 130 inputs the prompt message, the picture, and the first text into the large language model. The large language model recognizes the network link in the QR code according to the prompt message, the picture, and the first text, generates and outputs a first task instruction and a second text containing the information "Okay, about to log in to the XX website". The first task instruction is used to indicate an operation to log in to the XX website according to the recognized network link, the user account Jack, and the login password 123456 in the first text. The cloud server 130 executes the first task instruction, or forwards the first task instruction to the smart mobile terminal 120 or through the smart mobile terminal 120 to the smart glasses 110 for execution. At the same time, the cloud server 130 converts the second text into the corresponding voice and forwards the voice to the smart glasses 110 through the smart mobile terminal 120 for playback.
[0064] After the smart glasses 110 play the second voice, the user then takes 2 pictures through the smart glasses 110 and says the second voice command "What is the building in the picture?". The smart glasses 110 obtain the location information and send the location information, the 2 pictures, and the second voice command to the cloud server 130. The cloud server 130 converts the second voice command into the corresponding third text and performs semantic analysis on the third text to obtain a prompt message, where the prompt message includes a task request to recognize the pictures. Subsequently, the cloud server 130 inputs the prompt message, the 2 pictures, the location information, and the third text into the large language model. The large language model recognizes the buildings in the 2 pictures according to the prompt message, the location information, and the third text, and outputs a fourth text containing the introduction information of the building. The cloud server 130 converts the fourth text into the corresponding voice and forwards the voice to the smart glasses 110 through the smart mobile terminal 120 for playback. In practical applications, the 2 pictures may contain the same one or more buildings, or may contain multiple different buildings.
[0065] After the intelligent glasses 110 play the voice containing the introduction information of the building, the user then takes a video and utters the third voice "Share the video and only one of the two previously taken pictures that shows the building on XX website" through the intelligent glasses 110. The intelligent glasses 110 send the video and the third voice to the cloud server 130. The cloud server 130 obtains the historical chat record, converts the third voice into the corresponding fifth text, and performs semantic parsing on the fifth text to obtain a prompt message, where the prompt message includes a task request for the recognition operation. Subsequently, the cloud server 130 inputs the historical chat record, the prompt message, the video, and the fifth text into the large language model. The large language model identifies the picture that shows only the building among the two pictures as the target picture based on the historical chat record, the prompt message, the video, and the fifth text, and generates and outputs a second task instruction and a sixth text including "Okay, about to perform the sharing operation". The second task instruction is used to indicate sharing the video and the target picture to the above-mentioned XX website. The cloud server 130 executes the second task instruction, or forwards the second task instruction to the intelligent mobile terminal 120 or through the intelligent mobile terminal 120 to the intelligent glasses 110 for execution. Meanwhile, the cloud server 130 converts the sixth text into the corresponding voice and forwards the voice to the intelligent glasses 110 through the intelligent mobile terminal 120 for playback.
[0066] While performing the above series of operations, the cloud server 130 can also generate chat records in real time and save them in the historical chat record database for subsequent use (such as in the task request for the recognition operation).
[0067] To reduce the computing burden of a single server and improve the efficiency of data processing, in another embodiment, the above cloud server can be a distributed server cluster composed of multiple servers. The large language model (model server), historical chat record database (storage server), speech-to-text engine (speech-to-text server), text-to-speech engine (text-to-speech server), and prompt generator (prompt server) are respectively configured on multiple servers. The model server completes the above operations of obtaining the historical chat record, speech-to-text conversion, generating prompt information, obtaining the second text, text-to-speech conversion, etc. through data interaction with the storage server, speech-to-text server, text-to-speech server, and prompt server.
[0068] Optionally, in another embodiment of the present application, as Figure 4 shown, the cloud server 130 includes a model server 131, and the large language model is configured in the model server 131.
[0069] The intelligent mobile terminal 120 is further configured to convert the first voice into the first text, perform semantic parsing on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server 131.
[0070] The model server 131 is configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send it to the intelligent mobile terminal 120.
[0071] The intelligent mobile terminal 120 is further configured to convert the second text into the second voice.
[0072] That is to say, the intelligent mobile terminal 120 can also perform operations of first voice-to-text conversion, prompt information generation, and second text-to-voice conversion through a voice-to-text engine, a prompt generator, and a text-to-voice engine built in the intelligent mobile terminal 120.
[0073] Optionally, in another embodiment, the cloud server 130 further includes: a voice-to-text server 132, a prompt server 133, and a text-to-voice server 134. The intelligent mobile terminal 120 is further configured to: send the first voice to the voice-to-text server 132 to convert the first voice into the first text through the voice-to-text server 132; send the first text to the prompt server 133 to perform semantic parsing on the first text and generate the prompt information according to the parsed semantics through the prompt server 133; and send the second text to the text-to-voice server 134 to convert the second text into the second voice through the text-to-voice server 134.
[0074] That is to say, the intelligent mobile terminal 120 can also perform operations of first voice-to-text conversion, prompt information generation, and second text-to-voice conversion through data interaction with the voice-to-text server 132, the prompt server 133, and the text-to-voice server 134.
[0075] Optionally, in another embodiment, the cloud server 130 further includes a storage server 135;
[0076] The intelligent mobile terminal 120 is further configured to generate a real-time chat record, and save the real-time chat record in a database on the intelligent mobile terminal 120, or send it to the storage server 135 for saving;
[0077] The model server 131 is further configured to obtain the historical chat record sent by the intelligent mobile terminal 120, or obtain the historical chat record from the storage server 135.
[0078] Specifically, a database for storing historical chat records is configured on the intelligent mobile terminal 120 or the storage server 135. During the human-computer interaction (such as chatting) between the user and the large language model, the intelligent mobile terminal 120 can generate real-time chat records according to the received or sent data whenever receiving or sending data during the process of data interaction with each server of the intelligent glasses 110 and the cloud server 130, and save the generated real-time chat records in the database. The real-time chat records may include, but are not limited to, at least one of the received or sent images, voices, and texts. Further, the real-time chat records may also include at least one of the data interaction time, the identification information of both parties of the data interaction, and the user's identity identification information. Optionally, the real-time chat records may also be generated by the intelligent glasses 110 or the model server 131 and stored in the database on the storage server 135. Or, optionally, each server of the intelligent glasses 110, the intelligent mobile terminal 120, and the cloud server 130 may generate their respective corresponding real-time chat records and send them to the storage server 135, and then the storage server 135 sorts them out and stores them in the database.
[0079] Optionally, in another embodiment, the cloud server 130 includes a model server 131, a voice-to-text server 132, a prompt server 133, a text-to-voice server 134, and a storage server 135, and the large language model is configured in the model server 131.
[0080] The model server 131 is further configured to generate real-time chat records and send the real-time chat records to the storage server 135 for storage.
[0081] The model server 131 is further configured to: obtain the historical chat records from the storage server 135; send the first voice to the voice-to-text server 132 to convert the first voice into the first text through the voice-to-text server 132; send the first text to the prompt server 133 to perform semantic parsing on the first text and generate the prompt information according to the parsed semantics by the prompt server 133; and send the second text to the text-to-voice server 134 to convert the second text into the second voice through the text-to-voice server 134.
[0082] Optionally, in another embodiment, the cloud server 130 includes a model server 131, and the large language model is configured in the model server 131.
[0083] The intelligent glasses 110 are further configured to convert the first voice into the first text and send the first text and the image to the intelligent mobile terminal 120.
[0084] The intelligent mobile terminal 120 is further configured to perform semantic parsing on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server 131.
[0085] The model server 131 is configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send it to the intelligent mobile terminal 120.
[0086] The intelligent mobile terminal 120 is further configured to send the second text to the smart glasses 110.
[0087] The smart glasses 110 are further configured to convert the second text into the second voice.
[0088] Furthermore, the cloud server 130 further includes: a prompt server 133 and a storage server 135. The smart glasses 110 are further configured to send the first text to the prompt server 133, so that the prompt server 133 performs semantic parsing on the first text and generates the prompt information according to the parsed semantics. The model server 131 is further configured to obtain the historical chat record from the storage server 135.
[0089] Optionally, in another embodiment, the smart glasses 110 further include a positioning component. The smart glasses 110 are further configured to: obtain the positioning information through the positioning component, and send the first voice, the positioning information, and the image to the cloud server 130. At this time, the cloud server 130 is further configured to: through the large language model, obtain the second text according to the positioning information, the first text, the prompt information, and the image.
[0090] Specifically, the smart glasses 110 can directly communicate with the cloud server 130, that is, while obtaining the first voice, obtain the positioning information through its own positioning component, and send the first voice, the positioning information, and the image to the cloud server 130. After receiving the first voice, the positioning information, and the image, the cloud server 130 converts the first voice into the first text, performs semantic parsing on the first text and generates the prompt information according to the parsed semantics, obtains the second text according to the positioning information, the first text, the prompt information, and the image through the large language model, converts the second text into the second voice and returns it to the smart glasses 110 for playback. Alternatively, the cloud server 130 can also send the second voice to the intelligent mobile terminal 120 associated with the smart glasses 110 for playback through the intelligent mobile terminal 120 or forwarded to the smart glasses 110 through the intelligent mobile terminal 120.
[0091] Such direct communication between the smart glasses 110 and the cloud server 130 can reduce the data loss rate and improve the accuracy and efficiency of data processing.
[0092] Optionally, before positioning, the smart glasses 110 can first determine whether they are equipped with a positioning component according to the pre-stored device configuration information. If they have a positioning component, the positioning information is obtained through this positioning component. If they are not equipped with a positioning component, nearby connectable devices are searched via Bluetooth, a connection relationship is established with the searched connectable devices, and the positioning information is obtained through this connectable device. Or, if they are not equipped with a positioning component, it is determined whether a connection relationship has been established with the associated smart mobile terminal 120. If a connection relationship has been established, the positioning information is obtained from the smart mobile terminal 120. Further, if the smart glasses 110 are not equipped with a positioning component, there is no connected smart mobile terminal 120, and no connectable device is searched, the smart glasses 110 obtain the last sent positioning information associated with this user from the historical chat record database, use it as the positioning information and mark it, so that the cloud server 130 can verify this positioning information according to this mark. The verification method is, for example, if this user is bound to a companion, the cloud server 130 can request the smart glasses or smart mobile terminal of this companion to obtain the positioning information, and verify the marked positioning information according to the obtained positioning information. If the position deviation between the two is less than the preset value, then according to this large language model, according to this historical chat record, this positioning information, this first text, this prompt information, and this image, obtain this second text, otherwise, according to this large language model, according to this historical chat record, this first text, this prompt information, and this image, obtain this second text.
[0093] Optionally, before establishing a data connection with the connectable device, the smart glasses 110 can also verify the connectable device using a preset password, or only connect to the connectable devices in the preset security list to improve the security of data transmission.
[0094] Further, the smart glasses 110 are also used to send this first voice to the smart mobile terminal 120 to generate this prompt information through the smart mobile terminal 120 according to this first voice. The smart glasses 110 are also used to send this prompt information, this first voice, this positioning information, and this image to the cloud server 130. The cloud server 130 is also used to: obtain the historical chat record, and according to this large language model, according to this historical chat record, this positioning information, this first text, this prompt information, and this image, obtain this second text.
[0095] Optionally, in another embodiment, the smart glasses 110 further include at least one control button, which includes a physical button and / or a virtual button based on a touch sensor. The smart glasses 110 are further configured to: in response to an activation instruction, activate the chat function of the smart glasses 110, where the activation instruction comes from the smart mobile terminal 120 or a virtual assistant program built into the smart glasses 110; in response to a photographing instruction, capture the image through the camera, and the photographing instruction is triggered by the user through the photographing button in the control button; output a first prompt sound to prompt the user to ask a question; in response to a listening instruction, activate the microphone to start picking up the first voice, and the listening instruction is triggered based on the event that the user presses the chat button in the control button; in response to an end instruction, stop picking up the first voice and output a second prompt sound to prompt the user to end the voice pickup, and the end instruction is triggered based on the event that the user releases the chat button or the event that the idle time of the microphone exceeds a preset time.
[0096] Optionally, when sending the first voice or the first text, the smart glasses 110 or the smart mobile terminal 120 can also send one or more sensed data, such as the step count data of a pedometer, the user's posture, sensor data such as IMU (Inertial Measurement Unit) data, electronic compass direction data, touch sensor data, and environmental data such as temperature or humidity, obtained through built-in sensors to the large language model, so that the large language model can better respond to user queries. For example, if the user asks "How much longer do I need to walk to reach the location in the picture", the large language model will correctly answer the user's question about the remaining time by combining the location of the location recognized in the image and the step count data of the pedometer, such as "You still need to walk for half an hour to reach that location."
[0097] In this embodiment, by using the large language model, human-machine task interaction based on a captured image is realized on the smart glasses, thus enriching the functions of the smart glasses. And due to the scalability and self-creation of the large language model, the intelligence and interactivity of the smart glasses can be further improved. Further, since the smart glasses can perform positioning and data forwarding through nearby connectable devices when they have no wireless network or positioning components themselves, the success rate of data processing can be increased.
[0098] See Figure 5 , Figure 5 is a schematic internal structure diagram of the smart glasses provided by an embodiment of the present application. Figure 6 is a schematic external structure diagram of the smart glasses provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown in the figure. As Figure 5 and Figure 6As shown, the smart glasses 200 include: a frame 201, at least one temple 202, at least one microphone 203, at least one speaker 204, at least one processor 205, at least one memory 206, and at least one sensor 207. The at least one sensor 207 includes a camera 2071, and the camera 2071 can be disposed on the left side of the frame 201 (as Figure 6 shown), or can also be disposed on the right side of the frame 201, or can also be disposed in the exact middle of the frame 201, and the present application does not make specific limitations. Figure 5 and Figure 6 are only the best examples. In practical applications, the smart glasses 200 may also have fewer or more components than Figure 5 and Figure 6 shown.
[0099] The frame 201 can be, for example, a front frame with lenses (such as, sunglasses lenses, transparent lenses, or corrective lenses). The at least one temple 202 can include, for example, a left temple and a right temple.
[0100] The temple 202 is connected to the frame 201, and the processor 205 is electrically connected to the microphone 203, the speaker 204, the memory 205, and the sensor 207. The microphone 203, the speaker 204, the processor 205, the memory 206, and the sensor 207 are disposed on at least one temple 202 and / or the frame 201. Preferably, the temple 202 is detachably connected to the frame 201.
[0101] The processor 205 includes: a central processing unit (CPU) and a DSP (Digital Signal Processing). The DSP is used to process the voice data acquired by the microphone 203. The CPU is preferably an MCU (Microcontroller Unit).
[0102] The memory 206 is a non - transitory memory, which may specifically include: RAM (Random Access Memory) and flash memory components, and stores one or more programs executable by the processor 205. The one or more programs include multiple instructions. The multiple instructions are used for: obtaining sensing data through at least one sensor 207, where the at least one sensor 207 includes a camera 2071, and the sensing data includes an image captured by the camera 2071; obtaining the user's first voice through the microphone, where the first voice contains a question raised by the user; obtaining a second voice containing a reply to the question based on the first voice and the sensing data through a large - language model, where the large - language model is configured in the smart glasses or a smart mobile terminal or a cloud server; and playing the second voice through the speaker 204.
[0103] Among them, the large - language model can be a GAILLM model or an MLLM model or other models with similar functions.
[0104] Optionally, in other embodiments of the present application, the smart glasses 200 further include at least one control button 208 electrically connected to the processor 205. The control button 208 includes a physical button and / or a virtual button based on a touch sensor. The multiple instructions are further used for: in response to an activation instruction, activating the chat function of the smart glasses 200, where the activation instruction comes from an associated smart mobile terminal or a virtual assistant program built into the smart glasses 200; in response to a photographing instruction, taking the image through the camera 2071, where the photographing instruction is triggered by the user through a photographing button in the control button 208; outputting a first prompt sound to prompt the user to ask a question; in response to a listening instruction, activating the microphone 203 to start picking up the first voice, where the listening instruction is triggered based on an event that the user presses a chat button in the control button 208; in response to an end instruction, stopping picking up the first voice, outputting a second prompt sound to prompt the user to end the voice pickup, and performing the operation of obtaining a second voice containing a reply to the question based on the first voice and the sensing data through the large - language model, where the end instruction is triggered based on an event that the user releases the chat button or an event that the idle time of the microphone 203 exceeds a preset time.
[0105] Optionally, in other embodiments of the present application, the at least one sensor 207 further includes a position sensor 2072 ( Figure 5 and Figure 6 not shown in the figure), the sensing data further includes positioning data of the smart glasses 200, and the multiple instructions are further used for: obtaining the position information of the smart glasses 200 as the positioning data through the position sensor 2072. The function of the position sensor 2072 is the same as that of Figure 1 and Figure 4The positioning component in the illustrated embodiment of the control system is similar to a satellite positioning system, and specifically may but is not limited to be a positioning component based on GPS or Beidou satellites.
[0106] Optionally, in other embodiments of the present application, the plurality of instructions are further used for: obtaining the historical chat records of the user, and through the large language model, obtaining the second voice according to the first voice, the sensing data, and the historical chat records.
[0107] Optionally, in other embodiments of the present application, the smart glasses 200 further include a wireless communication component 208 electrically connected to the processor 205, and the plurality of instructions are further used for: generating real-time chat records and storing them in the memory, or sending the real-time chat records to a storage server for storage through the wireless communication component; obtaining the historical chat records from the memory or the storage server.
[0108] Specifically, the wireless communication component 208 includes a wireless signal transceiver and its surrounding circuits, and may be specifically disposed in the inner cavity of the spectacle frame 201 and / or at least one temple 202. The wireless signal transceiver may use but is not limited to at least one of the WiFi (Wireless Fidelity) protocol, NFC (Near Field Communication) protocol, ZigBee, UWB (UltraWideband), RFID (Radio Frequency Identification) protocol, and cellular mobile communication protocol (such as 3G / 4G / 5G, etc.) for data transmission.
[0109] Optionally, in other embodiments of the present application, the smart glasses 200 further include a Bluetooth component 209 electrically connected to the processor 205, and the plurality of instructions are further used for: obtaining the positioning data from the smart mobile terminal through the Bluetooth component 209; sending the first voice to a speech-to-text server through the wireless communication component 208 to convert the first voice into a first text by the speech-to-text server; sending the first text to a prompt server through the wireless communication component 208 to perform semantic parsing on the first text by the prompt server and generate the prompt information according to the parsed semantics; obtaining a second text including a reply to the question according to the prompt information, the first text, the sensing data, and the historical chat records through the large language model; and sending the second text to a text-to-speech server through the wireless communication component 208 to convert the second text into the second voice by the text-to-speech server.
[0110] Optionally, in other embodiments of the present application, the multiple instructions are further used to: send the prompt message, the first text, the sensing data, and the historical chat record to the smart mobile terminal through the Bluetooth component 209, so as to obtain a second text through the large language model on the smart mobile terminal, where the second text contains the reply to the question; or, send the prompt message, the first text, the sensing data, and the historical chat record to the model server through the wireless communication component 208, so as to obtain the second text through the large language model on the model server.
[0111] Optionally, in other embodiments of the present application, the smart glasses 200 further include other input devices electrically connected to the processor 205, such as: a power-on button, a touch display screen, and the like. Further, the touch display screen can also be used as an output device for displaying the real-time chat record or the user-specified historical chat record in the above other embodiments.
[0112] Optionally, in other embodiments of the present application, the smart glasses 200 further include an indicator light and / or a buzzer electrically connected to the processor 205. The multiple instructions are further used to output a prompt message through the indicator light and / or the buzzer. The prompt message is used to prompt the state of the smart glasses 200, and the state includes a working state and an idle state. The working state includes: a state of starting to pick up voice, a voice picking-up state, a state of ending voice picking-up, and a voice processing state. Among them, the indicator light can be an LED (Light Emitting Diode) light.
[0113] Specifically, in addition to using a beeping notification to indicate that the smart glasses are listening to the user's voice, to indicate that the voice has been sent to the large language model and is waiting for a response, the smart glasses can also use an LED to indicate whether the smart glasses are listening to the user's voice (for example, green), are waiting for a response from the large language model (red), or are completely idle (off). And after playing the voice (such as the second voice) obtained through the large language model, if the chat function of the smart glasses is turned off due to the microphone being idle for more than a preset duration, the LED is also turned off. At this time, if the user sees this LED turned off, they know that they can use the method of pressing the virtual button or the wake-up word to reactivate the smart glasses to listen to the user's voice again.
[0114] Optionally, in other embodiments of the present application, at least one sensor 207 further includes at least one device among an inertial measurement unit (IMU) sensor, a temperature sensor, a proximity sensor, a humidity sensor, an electronic compass, a timer, and a pedometer.
[0115] The multiple instructions are also used to obtain the sensing data of the at least one device in response to a sensing data acquisition request from the model server and send it to the model server, so that the model server can obtain the second text through the large language model based on the historical chat record, the location information, the first text, the prompt information, the image, and the sensing data. It can be understood that the text output by the large language model may also be other texts including the sensing data acquisition request. The model server generates a sensing data acquisition request according to the other text and sends it to the smart glasses 200, so that the smart glasses 200 can obtain the sensing data pointed to by the request according to the sensing data acquisition request. For example, if the user makes a task request of "how much longer do I have to walk to reach the location in the picture", the large language model will output a text containing a pedometer data acquisition instruction, and the model server sends a pedometer data acquisition request to the smart glasses 200 or a smart mobile terminal associated with the smart glasses 200 according to the pedometer data acquisition instruction, so as to obtain pedometer data from the smart glasses 200 or a smart mobile terminal associated with the smart glasses 200 and input it into the large language model. The large language model calculates the time for the user to reach the location according to the step count data of the pedometer and the position of the object recognized in the picture, and outputs a second text containing the information of "you still have to walk for half an hour to reach the location".
[0116] Optionally, in other embodiments of the present application, the smart glasses 200 further include a voice biometric module, which is used to identify the user by using the voiceprint of the user of the smart glasses 200 to enable the above-mentioned voice control function based on the smart glasses 200. Since only the owner of the smart glasses can wake up the smart glasses to use the functions provided by the large language model, the security of the device can be improved.
[0117] Furthermore, the smart glasses 200 further include a battery 210, which is used to provide power support for each electronic device of the above-mentioned smart glasses 200, such as the above-mentioned microphone 203, speaker 204, processor 205, memory 206, at least one sensor 207, and other components.
[0118] Each electronic component of the above-mentioned smart glasses can be connected through a bus.
[0119] It should be noted that the relationship between the components of the above smart glasses can be a substitution relationship or a superposition relationship. That is, all the components in the above embodiments can be installed on one smart glasses, or, alternatively, according to requirements, some of the above components can be selectively installed. When it is a substitution relationship, the smart glasses are also provided with a connection interface for external devices, and this connection interface can be, for example, at least one of a PS / 2 interface, a serial interface, a parallel interface, an IEEE1394 interface, a USB (Universal Serial Bus) interface, etc. The functions of the replaced components can be realized by external devices connected to this connection interface, such as external speakers, external sensors, etc.
[0120] For the details not described in this embodiment, reference can also be made to the relevant descriptions in the above Figures 1 to 4 and Figures 7 to 12 illustrated embodiments, which will not be elaborated here.
[0121] In this embodiment, by using a large language model, human-machine task interaction based on captured images is realized on the smart glasses, thus enriching the functions of the smart glasses. And due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses can be further improved. Further, since the smart glasses can also perform positioning and data forwarding by means of nearby connectable devices when they do not have a wireless network or a positioning component themselves, the success rate of data processing can be increased.
[0122] See Figure 7 , Figure 7 which is a flowchart of the implementation of a control method for a smart wearable device based on a large language model provided by an embodiment of the present application. This control method can be applied to a smart mobile terminal, such as Figure 1 and Figure 4 the smart mobile terminal 120 shown in the illustrated embodiments. The smart wearable device can implement this control method by means of data interaction with the smart mobile terminal 120. Among them, the smart wearable device can include, but is not limited to: a smart safety helmet, smart headphones, smart earrings, a smart watch, and smart glasses such as Figure 5 and Figure 6 shown. As Figure 7 shown, this control method includes the following steps:
[0123] S701: Receive the first voice and image sent by the smart wearable device through Bluetooth, where the first voice contains a question raised by the user;
[0124] S702: Convert the first voice into a first text, perform semantic parsing on the first text, and generate a prompt message according to the parsed semantics;
[0125] S703: Obtain a second text through a large language model based on the image, the first text, and the prompt information, where the second text contains a reply to the question, and the large language model is configured on the intelligent mobile terminal or the cloud server;
[0126] S704: Convert the second text into a second voice and send the second voice to the intelligent wearable device through the Bluetooth for playback.
[0127] Optionally, in other embodiments of the present application, the method further includes: obtaining location information and / or the user's historical chat records; and obtaining the second text through a large language model based on the image, the first text, the prompt information, and the location information and / or the historical chat records.
[0128] Optionally, in other embodiments of the present application, specifically converting the first voice into the first text may further include: sending the first voice to a voice-to-text server through a wireless network so that the voice-to-text server converts the first voice into the first text.
[0129] Semantically parsing the first text and generating prompt information according to the parsed semantics may specifically further include: sending the first text to a prompt server through the wireless network so that the prompt server semantically parses the first text and generates the prompt information according to the parsed semantics.
[0130] Converting the second text into a second voice may specifically further include: sending the second text to a text-to-voice server through the wireless network so that the text-to-voice server converts the second text into the second voice.
[0131] Optionally, in other embodiments of the present application, after converting the first voice into the first text, the method further includes: generating a real-time chat record according to the first text and the image, and storing the real-time chat record in the database of the intelligent mobile terminal or the storage server; displaying the real-time chat record through a display screen; after obtaining the second text, associating the second text with the real-time chat record in the database or the storage server; and displaying the reply in the second text through the display screen. The display screen may be a display screen built in or externally connected to the intelligent mobile terminal.
[0132] Obtaining the user's historical chat records may specifically further include: obtaining the historical chat records during the current chat from the database or the storage server.
[0133] Optionally, in other embodiments of the present application, the prompt information is used to prompt the large language model to perform at least one of the following tasks based on the image: image recognition, information sharing, making a call, navigation, searching the Internet, invoking a network server, and translation.
[0134] The above intelligent glasses control system and intelligent wearable device control method will be further described below in combination with multiple application examples.
[0135] Application Example 1
[0136] The large language model is configured in the server 130. The intelligent glasses 110 are responsible for picking up voices, playing voices, and taking pictures. The smartphone or smartwatch 120 is responsible for GPS positioning. The cloud server is responsible for voice-to-text conversion, text-to-voice conversion, obtaining historical chat records, generating prompt information, and obtaining a reply to the user's question through the large language model. Among them, the prompt information can also be generated by the smartphone or smartwatch 120.
[0137] As Figure 8 shown, after the user starts the chat function through a mobile application (APP) or a virtual assistant program (such as Siri / OK Google, etc.) running on the smartphone or smartwatch 120, if the user presses and holds the virtual button based on the touch sensor on the temple of the intelligent glasses 110, the intelligent glasses 110 call the built-in camera to take a picture. At the same time, a prompt sound is output through the built-in speaker to prompt the user that the microphone on the intelligent glasses 110 is ready to listen to the user's speech.
[0138] While pressing and holding the virtual button, the user asks a question such as "Where in Hong Kong was this photo taken" through voice. The intelligent glasses 110 pick up the first voice containing the question through the microphone. When the user releases the virtual button, the intelligent glasses 110 send the first voice and the taken photo to the mobile APP on the smartphone or smartwatch 120 via Bluetooth, and output another notification prompt sound to prompt the user that the task request has been sent.
[0139] Then, the mobile APP obtains the GPS positioning information, and sends the first voice, the photo, and the GPS positioning information to the cloud server 130. The cloud server 130 converts the first voice into a first text, semantically analyzes the first text, generates prompt information according to the parsed semantics, and obtains the user's historical chat records. The historical chat records, for example, can be the text corresponding to the user's voice and / or the sent photo when the user last chatted with the large language model before the current time point after the start of the current chat function. It can be understood that an exchange of questions and answers between the user and the large language model can be regarded as a chat.
[0140] After that, the cloud server 130 inputs the taken photo, the first text, the prompt information, the GPS positioning information, and the historical chat record into the large language model, and obtains a second text containing the user's question as the reply output by the large language model. The second text is converted into a second voice and sent to the smartphone or smartwatch 120. The smartphone or smartwatch 120 then sends the received second voice to the smart glasses 110 via Bluetooth so that the smart glasses 110 can play it using the built-in speaker.
[0141] Then, the user can continue to ask the next question until the user stops the chat function through the mobile APP.
[0142] Furthermore, the user can set the language they use through the mobile APP, or alternatively, the voice-to-text engine can automatically detect the user's language, and the prompt information generated by the cloud server 130 also includes information about the user's language.
[0143] Furthermore, the user can set the playback speed of the second voice through the mobile APP, for example: set to normal, 1.25 times faster or slower, 1.5 times faster or slower, 2 times faster or slower.
[0144] Furthermore, during the startup of the chat function, the smartphone or smartwatch 120 can generate a real-time chat record based on the received data, and display the real-time chat record to the user through the mobile APP, while associating the real-time chat record with the user account and storing it in the historical chat record database. Among them, the historical chat record database can be configured in the storage server or in the smartphone or smartwatch 120.
[0145] Furthermore, according to the user's operation on the mobile APP, the smartphone or smartwatch 120 can also export the chat record pointed to by the user operation as a file to a social media platform such as Facebook, etc.
[0146] Furthermore, the user can view the complete chat record or the chat record of a specified date or content through the mobile APP.
[0147] Furthermore, when the smart glasses 110 play the second voice, the user can slide the temple of the smart glasses 110 towards the ear direction to increase the volume, or slide the temple of the smart glasses 110 away from the ear direction to decrease the volume).
[0148] Application Example 2
[0149] The large language model is configured in the server 130. As Figure 9As shown, different from Application Example 1, in Application Example 2, when the user stops speaking and is idle for a period of time (i.e., the idle time of the microphone exceeds the preset time), the smart glasses 110 stop picking up sound and send the first voice and the taken photo to the smartphone or smartwatch 120.
[0150] Application Example 3
[0151] The large language model is configured in the server 130. As Figure 10 shown, different from Application Example 1, in Application Example 3, the smart glasses 110 are responsible for picking up sound, playing sound, taking photos, speech-to-text conversion, and text-to-speech conversion, the smartphone or smartwatch 120 is responsible for GPS positioning, and the cloud server 130 is responsible for generating prompt information, obtaining historical chat records, and getting answers to the user's questions through the large language model.
[0152] Application Example 4
[0153] The large language model is configured in the server 130. As Figure 11 shown, different from Application Example 1, in Application Example 4, the smart glasses 110 are responsible for picking up sound, playing sound, and taking photos, the smartphone or smartwatch 120 is responsible for GPS positioning, speech-to-text conversion, and text-to-speech conversion, and the cloud server 130 is responsible for generating prompt information, obtaining historical chat records, and getting answers to the user's questions through the large language model.
[0154] Application Example 5
[0155] The large language model is configured in the server 130. As Figure 12 shown, different from Application Example 1, in Application Example 5, the smart glasses 120 communicate directly with the cloud server 130. The smart glasses 110 are responsible for picking up sound, playing sound, GPS positioning, and taking photos, and the cloud server 130 is responsible for speech-to-text conversion, text-to-speech conversion, generating prompt information, obtaining historical chat records, and getting answers to the user's questions through the large language model.
[0156] For the details not described in this embodiment, reference can also be made to the relevant descriptions in the above Figures 1 to 6 shown embodiment, which will not be elaborated here.
[0157] In this embodiment, by using the large language model, human-machine task interaction based on the captured image is realized on the smart glasses, thus enriching the functions of the smart glasses. And due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses can be further improved.
[0158] The embodiments of the present application also provide a non-transitory computer-readable storage medium, which can be disposed in the smart glasses or smart wearable devices in the above embodiments. The non-transitory computer-readable storage medium can be the memory 206 in the Figure 5 embodiment shown above. A computer program is stored on the computer-readable storage medium. When the program is executed by a processor, it implements the control method of the smart wearable device based on the large language model described in the above embodiments. Further, the computer-readable storage medium can also be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk, or an optical disc that can store program codes.
[0159] In several embodiments provided by the present application, it should be understood that the disclosed smart glasses, control system, and smart wearable device control method can be implemented in other ways. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed direct or direct connection or communication connection between each other can be an indirect connection or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.
[0160] It should be noted that for the above embodiments of the smart wearable device control method, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0161] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] The above is the description of the smart glasses, control system, and smart wearable device control method provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An intelligent glasses control system based on a large language model, characterized in that, The system includes: smart glasses, a smart mobile terminal, and a cloud server, wherein the smart glasses include a microphone, a speaker, a camera, and a Bluetooth component, and the large language model is configured in the cloud server; The smart glasses are used to capture images through the camera, obtain the user's first voice through the microphone, and send the image and the first voice to the smart mobile terminal through the Bluetooth component, wherein the first voice contains the question raised by the user; The smart mobile terminal is used to send the first voice and the image to the cloud server; The cloud server is used to: convert the first voice into a first text; perform semantic analysis on the first text and generate a prompt message according to the parsed semantics; through the large language model, obtain a second text according to the first text, the prompt message, and the image, wherein the second text contains the answer to the question; convert the second text into a second voice; and send the second voice to the smart mobile terminal; The smart mobile terminal is further used to send the second voice to the smart glasses; The smart glasses are further used to receive the second voice through the Bluetooth component and play the second voice through the speaker.
2. The smart glasses control system according to claim 1, wherein The smart mobile terminal is further used to obtain positioning information and send the positioning information, the first voice, and the image to the cloud server, wherein the positioning information is obtained by the smart mobile terminal through a built-in satellite positioning system, or is obtained by the smart glasses and sent to the smart mobile terminal; The cloud server is further used to obtain the second text through the large language model according to the positioning information, the first text, the prompt message, and the image.
3. The smart glasses control system according to claim 1, wherein The cloud server is further used to obtain historical chat records and obtain the second text through the large language model according to the historical chat records, the first text, the prompt message, and the image.
4. The smart glasses control system according to claim 3, wherein The smart mobile terminal is further used to obtain positioning information and send the positioning information, the first voice, and the image to the cloud server, wherein the positioning information is obtained by the smart mobile terminal through a built-in satellite positioning system, or is obtained by the smart glasses and sent to the smart mobile terminal; The cloud server is further used to obtain the historical chat records and obtain the second text through the large language model according to the historical chat records, the positioning information, the first text, the prompt message, and the image.
5. The intelligent glasses control system according to claim 4, characterized in that, The cloud server includes a model server, and the large language model is configured in the model server; The intelligent mobile terminal is further configured to convert the first voice into the first text, perform semantic analysis on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server; The model server is configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send it to the intelligent mobile terminal; The intelligent mobile terminal is further configured to convert the second text into the second voice.
6. The intelligent glasses control system according to claim 5, characterized in that The cloud server further includes: a voice-to-text server, a prompt server, and a text-to-voice server. The intelligent mobile terminal is further configured to: Send the first voice to the voice-to-text server to convert the first voice into the first text through the voice-to-text server; Send the first text to the prompt server to perform semantic analysis on the first text through the prompt server and generate the prompt information according to the parsed semantics; and Send the second text to the text-to-voice server to convert the second text into the second voice through the text-to-voice server.
7. The intelligent glasses control system according to claim 6, wherein, The cloud server further includes a storage server; The intelligent mobile terminal is further configured to generate a real-time chat record, and save the real-time chat record in a database on the intelligent mobile terminal, or send it to the storage server for saving; The model server is further configured to obtain the historical chat record sent by the intelligent mobile terminal, or obtain the historical chat record from the storage server.
8. The intelligent glasses control system according to claim 4, characterized in that, The cloud server further includes: a model server, a voice-to-text server, a prompt server, a text-to-voice server, and a storage server. The large language model is configured in the model server; The model server is further configured to generate a real-time chat record and send the real-time chat record to the storage server for saving; The model server is further configured to: Obtain the historical chat record from the storage server; Send the first voice to the voice-to-text server to convert the first voice into the first text through the voice-to-text server; Send the first text to the prompt server to perform semantic analysis on the first text through the prompt server and generate the prompt information according to the parsed semantics; and Send the second text to the text-to-voice server to convert the second text into the second voice through the text-to-voice server.
9. The intelligent glasses control system according to claim 4, characterized in that The cloud server includes a model server, and the large language model is configured in the model server; The intelligent glasses are further configured to convert the first voice into the first text, and send the first text and the image to the intelligent mobile terminal; The intelligent mobile terminal is further configured to perform semantic parsing on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server; The model server is configured to obtain the historical chat record, and use the large language model to obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send the second text to the intelligent mobile terminal; The intelligent mobile terminal is further configured to send the second text to the intelligent glasses; The intelligent glasses are further configured to convert the second text into the second voice.
10. The intelligent glasses control system according to claim 9, characterized in that The cloud server further includes: a prompt server and a storage server; The intelligent glasses are further configured to send the first text to the prompt server, so that the prompt server performs semantic parsing on the first text and generates the prompt information according to the parsed semantics; The model server is further configured to obtain the historical chat record from the storage server.
11. The intelligent glasses control system according to claim 1, wherein The intelligent glasses further include a positioning component, and the intelligent glasses are further configured to: obtain positioning information through the positioning component, and send the first voice, the positioning information, and the image to the cloud server; The cloud server is further configured to use the large language model to obtain the second text according to the positioning information, the first text, the prompt information, and the image.
12. The intelligent glasses control system according to claim 11, wherein The intelligent glasses are further configured to send the first voice to the intelligent mobile terminal, so that the intelligent mobile terminal generates the prompt information according to the first voice; The intelligent glasses are further configured to send the prompt information, the first voice, the positioning information, and the image to the cloud server; The cloud server is further configured to obtain the historical chat record, and use the large language model to obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image.
13. The intelligent glasses control system according to claim 1, characterized in that, The intelligent glasses further include at least one control button, and the control button includes a physical button and / or a virtual button based on a touch sensor. The intelligent glasses are further configured to: Respond to an activation instruction to activate the chat function of the intelligent glasses, where the activation instruction comes from the intelligent mobile terminal or a virtual assistant program built in the intelligent glasses; Respond to a photographing instruction to capture the image through the camera, where the photographing instruction is triggered by the user through the photographing button in the control button; Output a first prompt sound to prompt the user to ask a question; Respond to a listening instruction to activate the microphone to start picking up the first voice, where the listening instruction is triggered based on an event that the user presses the chat button in the control button; In response to an end instruction, stop picking up the first voice, and output a second prompt tone to prompt the user to end voice pickup. The end instruction is triggered based on an event that the user releases the chat button or an event that the idle time of the microphone exceeds a preset time.
14. The intelligent glasses control system according to claim 1, wherein The large language model includes a generative artificial intelligence large language model or a multimodal large language model.
15. An intelligent glasses based on large language model, characterized in that, It includes: a spectacle frame, at least one temple, a microphone, a speaker, at least one sensor, a processor, and a memory; The at least one temple is connected to the spectacle frame, and the processor is electrically connected to the microphone, the speaker, the at least one sensor, and the memory; One or more programs executable by the processor are stored in the memory. The one or more programs include a plurality of instructions for: obtaining sensing data through the at least one sensor, where the at least one sensor includes a camera, and the sensing data includes an image captured by the camera; obtaining a first voice of the user through the microphone, where the first voice contains a question raised by the user; obtaining, through a large language model, a second voice containing a reply to the question based on the first voice and the sensing data, where the large language model is configured in the smart glasses or a smart mobile terminal or a cloud server; playing the second voice through the speaker.
16. The smart glasses according to claim 15, characterized in that, The smart glasses further include at least one control button electrically connected to the processor. The control button includes a physical button and / or a virtual button based on a touch sensor. The plurality of instructions are further used for: responding to an activation instruction to activate the chat function of the smart glasses, where the activation instruction comes from the smart mobile terminal or a virtual assistant program built in the smart glasses; responding to a photographing instruction to capture the image through the camera, where the photographing instruction is triggered by the user through a photographing button in the control button; outputting a first prompt tone to prompt the user to ask a question; responding to a listening instruction to activate the microphone to start picking up the first voice, where the listening instruction is triggered based on an event that the user presses the chat button in the control button; responding to an end instruction to stop picking up the first voice, outputting a second prompt tone to prompt the user to end voice pickup, and performing the operation of obtaining, through the large language model, a second voice containing a reply to the question based on the first voice and the sensing data, where the end instruction is triggered based on an event that the user releases the chat button or an event that the idle time of the microphone exceeds a preset time.
17. The smart glasses according to claim 16, characterized in that, The at least one sensor further includes a position sensor, and the sensing data further includes positioning data of the smart glasses. The plurality of instructions are further used for: obtaining the position information of the smart glasses as the positioning data through the position sensor.
18. The smart glasses according to claim 17, characterized in that, The plurality of instructions are further used for: obtaining the historical chat record of the user, and obtaining the second voice through the large language model based on the first voice, the sensing data, and the historical chat record.
19. The smart glasses according to claim 18, characterized in that, The smart glasses further include a wireless communication component electrically connected to the processor, and the plurality of instructions are further configured to: Generate a real-time chat record and save it in the memory, or send the real-time chat record to a storage server for saving through the wireless communication component; Obtain the historical chat record from the memory or the storage server.
20. The smart glasses according to claim 19, characterized in that, The smart glasses further include a Bluetooth component electrically connected to the processor, and the plurality of instructions are further configured to: Obtain the positioning data from the smart mobile terminal through the Bluetooth component; Send the first voice to a speech-to-text server through the wireless communication component, so as to convert the first voice into a first text by the speech-to-text server; Send the first text to a prompt server through the wireless communication component, so as to perform semantic parsing on the first text by the prompt server and generate the prompt information according to the parsed semantics; Obtain a second text containing the reply to the question according to the prompt information, the first text, the sensing data, and the historical chat record through the large language model; And Send the second text to a text-to-speech server through the wireless communication component, so as to convert the second text into the second voice by the text-to-speech server.
21. The smart glasses according to claim 20, characterized in that, The plurality of instructions are further configured to: Send the prompt information, the first text, the sensing data, and the historical chat record to the smart mobile terminal through the Bluetooth component, so as to obtain a second text containing the reply to the question through the large language model on the smart mobile terminal; Or, Send the prompt information, the first text, the sensing data, and the historical chat record to a model server through the wireless communication component, so as to obtain the second text through the large language model on the model server.
22. A control method for an intelligent wearable device based on a large language model, characterized in that, Applied to a smart mobile terminal, the method includes: Receive a first voice and an image sent by a smart wearable device through Bluetooth, where the first voice includes a question raised by a user; Convert the first voice into a first text, perform semantic parsing on the first text, and generate prompt information according to the parsed semantics; Obtain a second text containing the reply to the question according to the image, the first text, and the prompt information through a large language model, where the large language model is configured on the smart mobile terminal or a cloud server; Convert the second text into a second voice, and send the second voice to the smart wearable device through the Bluetooth for playing.
23. The control method according to claim 22, wherein The method further includes: Obtain positioning information and / or the user's historical chat record; Obtain the second text according to the image, the first text, the prompt information, and the positioning information and / or the historical chat record through a large language model.
24. The control method according to claim 23, wherein The converting the first voice into a first text includes: Send the first voice to a voice-to-text server via a wireless network to convert the first voice into the first text by the voice-to-text server; The semantic parsing of the first text and generating a prompt message according to the parsed semantics includes: Send the first text to a prompt server via the wireless network to perform semantic parsing on the first text by the prompt server and generate the prompt message according to the parsed semantics; The conversion of the second text into a second voice includes: Send the second text to a text-to-voice server via the wireless network to convert the second text into the second voice by the text-to-voice server.
25. The control method according to claim 23, wherein After converting the first voice into the first text, the method further includes: Generate a real-time chat record according to the first text and the image, and store the real-time chat record in a database or a storage server of the intelligent mobile terminal; Display the real-time chat record through a display screen; After obtaining the second text, associate the second text with the real-time chat record in the database or the storage server; Display the reply in the second text through the display screen; The obtaining of the user's historical chat record includes: Obtain the historical chat record during the current chat from the database or the storage server.
26. The control method according to claim 23, wherein The prompt message is used to prompt the large language model to perform at least one of the following tasks based on the image: image recognition, information sharing, making a call, navigation, searching the Internet, calling a network server, and translation.
Citation Information
Cited By
Intelligent glasses control method and device, computing equipment and system
CN120832683A