Large language model-based smart glasses control system, method, and smart glasses

Through the smart glasses control system with a large language model, the coordinated work of smart glasses with cloud servers and mobile terminals is realized, solving the problem of single function of smart glasses, improving intelligence and interactivity, and expanding functional applications.

WO2025148664A1PCT designated stage expired Publication Date: 2025-07-17SOLOS TECH SHENZHEN LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141261
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-12-22
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing smart glasses have single functions, low intelligence, expensive, and lack diverse interactions and functions.

Method used

The smart glasses control system based on the large language model is adopted, and the image and voice conversion and processing are realized through the collaborative work of smart glasses, smart mobile terminals and cloud servers. The large language model is used to provide rich replies and functions, such as image recognition, translation, navigation, etc.

Benefits of technology

It enriches the functions of smart glasses, improves its intelligence and interactivity, enhances the user experience, and reduces the dependence on data processing on wireless networks or positioning components alone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141261_17072025_PF_FP_ABST
    Figure CN2024141261_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A large language model-based smart glasses control system (100), a control method, and smart glasses (110). The system (100) comprises smart glasses (110), a smart mobile terminal (120), and a cloud server (130). The smart glasses (110) are used for capturing an image by means of a camera, acquiring a first voice of a user, and forwarding the image and the first voice to the cloud server (130) by means of the smart mobile terminal (120), wherein the first voice comprises a question of the user. The cloud server (130) is used for converting the first voice into a first text, generating prompt information, obtaining, by means of a large language model and on the basis of the first text, the prompt information and the image, a second text containing a reply to the question, converting the second text into a second voice and forwarding the second voice to the smart glasses (110) by means of the smart mobile terminal (120) for playback. The intelligence and interactivity of the smart glasses (110) are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Smart glasses control system, method and smart glasses based on large language model

[0001] This application claims priority to Chinese patent application number CN 2024100347951, entitled “Smart glasses control system, method and smart glasses based on large language model”, filed with the Patent Office of the State Intellectual Property Office of China on January 9, 2024, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of smart glasses technology, and in particular to a smart glasses control system based on a large language model, a smart glasses, and a smart wearable device control method. Background Art

[0003] With the development of computer technology, smart glasses are becoming more and more popular. However, existing smart glasses are expensive and, in addition to their own functions as smart glasses, usually only have the functions of listening to music and making or receiving calls. The functions are relatively simple and the degree of intelligence is relatively low. Technical issues

[0004] The embodiments of the present application provide a smart glasses control system, smart glasses, and smart wearable device control method based on a large language model, which are used to enrich the functions of smart glasses and improve the intelligence and interactivity of smart glasses. Technical Solutions

[0005] In one aspect, an embodiment of the present application provides a smart glasses control system based on a large language model, the system comprising: smart glasses, a smart mobile terminal, and a cloud server, wherein the smart glasses include a microphone, a speaker, a camera, and a Bluetooth component, and the cloud server is configured with the large language model;

[0006] The smart glasses are configured to capture an image using the camera, acquire a first voice of the user using the microphone, and send the image and the first voice to the smart mobile terminal via the Bluetooth component, wherein the first voice includes a question raised by the user;

[0007] The smart mobile terminal is configured to send the first voice and the image to the cloud server;

[0008] The cloud server is configured to: convert the first speech into a first text; perform semantic analysis on the first text and generate prompt information based on the analyzed semantics; obtain a second text based on the first text, the prompt information, and the image using the large language model, wherein the second text includes an answer to the question; convert the second text into a second speech; and send the second speech to the smart mobile terminal;

[0009] The smart mobile terminal is further configured to send the second voice to the smart glasses;

[0010] The smart glasses are further configured to receive the second voice via the Bluetooth component and play the second voice via the speaker.

[0011] In one aspect, an embodiment of the present application further provides smart glasses based on a large language model, comprising: a frame, at least one temple, a microphone, a speaker, at least one sensor, a processor, and a memory;

[0012] The at least one temple is connected to the frame, and the processor is electrically connected to the microphone, the speaker, the at least one sensor, and the memory;

[0013] The memory stores one or more programs executable by the processor, wherein the one or more programs include multiple instructions, and the multiple instructions are used to:

[0014] Acquiring sensory data through the at least one sensor, wherein the at least one sensor includes a camera, and the sensory data includes an image captured by the camera;

[0015] Acquiring a first voice of the user through the microphone, where the first voice includes a question raised by the user;

[0016] Obtaining, by a large language model, a second speech including an answer to the question based on the first speech and the sensor data, wherein the large language model is configured in the smart glasses, the smart mobile terminal, or the cloud server;

[0017] Play the second voice through the speaker.

[0018] In one aspect, an embodiment of the present application further provides a method for controlling a smart wearable device based on a large language model, which is applied to a smart mobile terminal. The method includes:

[0019] receiving a first voice and an image sent by the smart wearable device via Bluetooth, wherein the first voice includes a question raised by the user;

[0020] Converting the first speech into a first text, performing semantic analysis on the first text, and generating prompt information according to the analyzed semantics;

[0021] Obtaining a second text based on the image, the first text, and the prompt information using a large language model, wherein the second text includes an answer to the question, and the large language model is configured on the smart mobile terminal or the cloud server;

[0022] The second text is converted into a second voice, and the second voice is sent to the smart wearable device via the Bluetooth for playback. Beneficial effects

[0023] In various embodiments of the present application, the control system utilizes a large language model to implement image-based human-computer task interaction in smart glasses or smart wearable devices, thereby enriching the functionality of the smart glasses. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses or smart wearable devices can be further enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0025] FIG1 is a schematic diagram of the structure of a smart glasses control system based on a large language model according to an embodiment of the present application;

[0026] Figures 2 and 3 are schematic diagrams of application scenarios of the control system shown in Figure 1;

[0027] FIG4 is a schematic diagram of the structure of a smart glasses control system based on a large language model provided by another embodiment of the present application;

[0028] FIG5 is a schematic diagram of the internal structure of smart glasses provided in one embodiment of the present application;

[0029] FIG6 is a schematic diagram of the external structure of smart glasses provided in one embodiment of the present application;

[0030] FIG7 is a flowchart of an implementation method of a smart wearable device control method based on a large language model according to an embodiment of the present application;

[0031] FIG8 is a schematic diagram of Application Example 1 of the method shown in FIG7 ;

[0032] FIG9 is a schematic diagram of Application Example 2 of the method shown in FIG7 ;

[0033] FIG10 is a schematic diagram of Application Example 3 of the method shown in FIG7 ;

[0034] FIG11 is a schematic diagram of Application Example 4 of the method shown in FIG7 ;

[0035] FIG12 is a schematic diagram of Application Example 5 of the method shown in FIG7 . Modes for Carrying Out the Invention

[0036] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] Hereinafter, the terms "including", "having" and their cognates, which may be used in various embodiments of the present invention, are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or the possibility of adding one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0038] Furthermore, the terms “first,” “second,” “third,” etc., are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.

[0039] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present invention pertain. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized or overly formal meaning unless clearly defined in the various embodiments of the present invention.

[0040] Referring to Figure 1 , which is a schematic diagram of the structure of a natural language command control system based on a generative artificial intelligence large language model according to one embodiment of the present application, the control system 100 includes: smart glasses 110 , a smart mobile terminal 120 , and a cloud server 130 .

[0041] The cloud server 130 may be a single server or a distributed server cluster consisting of multiple servers.

[0042] The smart glasses 110 may be open smart glasses, including components such as a microphone, a speaker, a camera, and a Bluetooth component. The specific structure of the smart glasses 110 can be seen in the relevant description of the embodiments shown in Figures 5 and 6 below.

[0043] The smart mobile terminal 120 may include, but is not limited to, a cellular phone, a smartphone, other wireless communication devices, a personal digital assistant, an audio player, other media players, a music recorder, a video recorder, a camera, other media recorders, a smart radio, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a Moving Picture Experts Group (MPEG-1 or MPEG-2) Audio Layer 3 (MP3) player, a digital camera, and a smart wearable device (such as a smart watch, a smart bracelet, etc.). The smart mobile terminal 120 may also be installed with an Android, iOS, or other operating system.

[0044] The smart glasses 110 are configured to capture an image using a camera, capture a first voice message from a user using a microphone, and transmit the image and the first voice message to the smart mobile terminal 120 via a Bluetooth component. The first voice message includes a question posed by the user. The image can be a static picture, a dynamic video, or an image, or a combination of both. The number of images can be one or more.

[0045] The smart mobile terminal 120 is configured to send the first voice and the image to the cloud server 130 .

[0046] The cloud server 130 is configured to: convert the first speech into a first text; perform semantic analysis on the first text and generate prompt information based on the analyzed semantics; obtain a second text based on the first text, the prompt information, and the image using a large language model (LLM), wherein the second text includes an answer to the question; convert the second text into a second speech; and send the second speech to the smart mobile terminal 120.

[0047] The smart mobile terminal 120 is further configured to send the second voice to the smart glasses 110 .

[0048] The smart glasses 110 are further configured to receive the second voice via the Bluetooth component and play the second voice via the speaker.

[0049] The large language model may be configured on the cloud server 130. The large language model may include, but is not limited to, a Generative Artificial Intelligence Large Language Model (GAILLM) or a Multimodal Large Language Model (MLLM).

[0050] Examples of the generative AI large language model include, but are not limited to, Open AI's ChatGPT, Google's Bard, and other models with similar functionality. Examples of the multimodal large language model include, but are not limited to, BLIP-2, LLaVA, MiniGPT-4, mPLUG-Owl, LLaMA-Adapter-v2, Otter, Multimodal-GPT, InstructBLIP, VisualGLM-6B, PandaGPT, LaVIN, and other models with similar functionality. The large language model can include multiple components, each trained based on different samples, to answer different task requests based on the image, including, but not limited to, image recognition, information sharing, making phone calls, navigation, translation, sending messages, sending emails, searching the web, invoking network services, and invoking other third-party software development kits (SDKs) to perform tasks provided by the SDK. Specifically, user questions such as "Where is this photo in Hong Kong?", "What am I looking at?", "Which direction should I go?", "Log in to a website," or "Can the text in this photo be translated?"

[0051] Taking translation as an example, when a user shoots a video (or one or more pictures) through the smart glasses 110 and speaks the first voice "Translate the English in the video into Chinese", the smart glasses 110 sends the video and the first voice to the smart mobile terminal 120 through the Bluetooth component, so that the video and the first voice are forwarded to the cloud server 130 through the smart mobile terminal 120.

[0052] After receiving the video and the first speech, the cloud server 130 converts the first speech into a first text, performs semantic analysis on the first text, generates prompt information based on the analyzed semantics, and then inputs the first text, the prompt information, and the video into the large language model. The prompt information includes a translation task request.

[0053] Based on the first text and the prompt information, the large language model extracts text from the video (e.g., "The next station is Guangzhou South Railway Station"), translates the text using English as the source language and Chinese as the target language, and finally outputs a second text containing the translated content (e.g., "The next station is Guangzhou South Station"). Optionally, the large language model can also automatically identify the source language when performing the text extraction operation in the video.

[0054] The cloud server 130 converts the second text output by the large language model into a second voice and forwards it to the smart glasses 110 through the smart mobile terminal 120 for playback.

[0055] Optionally, in another embodiment of the present application, the smart mobile terminal 120 is further configured to obtain positioning information and send the positioning information, the first voice, and the image to the cloud server 130, wherein the positioning information is obtained by the smart mobile terminal 120 through a built-in satellite positioning system, or is obtained by the smart glasses 110 and sent to the smart mobile terminal 120. The cloud server 130 is further configured to obtain the second text based on the positioning information, the first text, the prompt information, and the image using the large language model. The satellite positioning system may be a GPS (Global Positioning System), or the satellite positioning system may also utilize Beidou satellites or other similar satellites for positioning.

[0056] Optionally, in another embodiment of the present application, the cloud server 130 is also used to obtain historical chat records, and obtain the second text based on the historical chat records, the first text, the prompt information and the image through the large language model.

[0057] Optionally, in another embodiment of the present application, the smart mobile terminal 120 is further configured to obtain positioning information and send the positioning information, the first voice, and the image to the cloud server 130, wherein the positioning information is obtained by the smart mobile terminal 120 through a built-in satellite positioning system, or is obtained by the smart glasses 110 and sent to the smart mobile terminal 120;

[0058] The cloud server 130 is further configured to obtain historical chat records and obtain the second text based on the historical chat records, the location information, the first text, the prompt information, and the image through the large language model.

[0059] As shown in Figure 2, in one practical application example, when a user discovers a new building while shopping, they can first take a photo using the camera on their smart glasses 110 and then ask a question such as "What building is this?" to the smart glasses 110. The smart glasses 110 capture the user's voice as a first voice message through their built-in microphone and then send the captured image and the first voice message to the smartphone 120 via Bluetooth.

[0060] The smart phone 120 obtains positioning information through a built-in positioning component, and then sends the positioning information, the received image and the first voice to the cloud server 130. The positioning component can be, but is not limited to, a positioning component based on GPS or Beidou satellites.

[0061] The cloud server 130 obtains the user's historical chat records, wherein the historical chat records include all texts and / or images interacted between the smart glasses 110, the smart phone 120 and the cloud server 130 when the user chats with the large language model through the smart glasses 110 during the activation period of the chat function of the smart glasses 110 or during a preset period.

[0062] At the same time, the cloud server 130 converts the first speech into first text, performs semantic analysis on the first text, and generates a prompt message based on the analyzed semantics, wherein the prompt message includes the analyzed semantics. As shown in FIG3 , in other application examples, based on the specific content of the analyzed semantics, in addition to image recognition, the prompt message can also be used to prompt the large language model to perform at least one of the following tasks based on the image: information sharing, making a phone call, navigation, translation, sending a message, sending an email, searching the web, invoking a web service, and invoking a third-party SDK to perform tasks provided by the SDK. The control system uses the large language model through the cloud server 130 to process and interpret the user's voice commands in natural language, and then performs the tasks specified by the voice context like a personal assistant. These tasks can include asking the model to find the name of a building after taking a photo with a camera, making a phone call, sending a message, sending an email, searching for interests, invoking a web service to retrieve / publish data, location, or navigation, and invoking a third-party SDK to perform tasks provided by the SDK. For example, after a user takes a photo of a building and then asks for a task such as "I want to know the name of this building," the large language model will respond to the user's task request.

[0063] Then, the cloud server 130 inputs the historical chat record, the positioning information, the first text, the prompt information and the image into the large language model, and obtains the second text output by the large language model. The second text contains the answer to the question. The cloud server 130 converts the second text into a second voice, and then sends the second voice to the smartphone 120, so that the smartphone 120 sends the second voice to the smart glasses 110 via Bluetooth for playback. The large language model uses the positioning information according to the prompt information to identify the building in the image, obtains the name of the building, and outputs a second text containing the name. Furthermore, the second text may also include an introduction to the building, such as its origin, structure, allusions, etc.

[0064] In another application example, a user can first use smart glasses 110 to capture a picture containing a QR code and speak a first voice message, "Log in to the website in the picture as Jack, using password 123456." Smart glasses 110 then send the picture and the first voice message to cloud server 130 via smart mobile terminal 120. Cloud server 130 converts the first voice message into corresponding first text and performs semantic parsing on the first text to obtain prompt information, where the prompt information includes a task request to recognize the picture. Cloud server 130 then inputs the prompt information, the picture, and the first text into a large language model. Based on the prompt information, the picture, and the first text, the large language model identifies the web link in the QR code and generates and outputs a first task instruction and a second text message containing the message "OK, about to log in to the XX website." The first task instruction instructs the user to log in to the XX website based on the identified web link and the user account Jack and login password 123456 in the first text message. The cloud server 130 executes the first task instruction, or forwards the first task instruction to the smart mobile terminal 120 or forwards it to the smart glasses 110 through the smart mobile terminal 120 for execution. At the same time, the cloud server 130 converts the second text into corresponding speech and forwards the speech to the smart glasses 110 through the smart mobile terminal 120 for playback.

[0065] After the smart glasses 110 play the second voice message, the user then uses them to take two more pictures and speaks a second voice message, "What are the buildings in the pictures?" The smart glasses 110 obtain location information and send the location information, the two pictures, and the second voice message to the cloud server 130. The cloud server 130 converts the second voice message into a corresponding third text and performs semantic parsing on the third text to obtain a prompt message, which includes a task request to identify the pictures. The cloud server 130 then inputs the prompt message, the two pictures, the location information, and the third text message into the large language model. Based on the prompt message, the location information, and the third text message, the large language model identifies the buildings in the two pictures and outputs a fourth text message containing information about the buildings. The cloud server 130 converts the fourth text message into a corresponding voice message and forwards the voice message to the smart glasses 110 via the smart mobile terminal 120 for playback. In practice, the two pictures may contain one or more identical buildings, or they may contain multiple different buildings.

[0066] After smart glasses 110 play a voice message introducing the building, the user then shoots a video and uses smart glasses 110 to speak a third voice message: "Share this video and the one of the two pictures you took previously that only shows the building on the XX website." Smart glasses 110 transmit the video and the third voice message to cloud server 130. Cloud server 130 retrieves the historical chat log, converts the third voice message into a corresponding fifth text, and performs semantic parsing on the fifth text to obtain a prompt message, which includes a task request for a recognition operation. Cloud server 130 then inputs the historical chat log, the prompt message, the video, and the fifth text into the large language model. Based on the historical chat log, the prompt message, the video, and the fifth text, the large language model identifies the picture of the two pictures that only shows the building as the target image, and generates and outputs a second task instruction and a sixth text message containing the message "OK, sharing will begin." The second task instruction instructs the user to share the video and the target image on the aforementioned XX website. The cloud server 130 executes the second task instruction, or forwards the second task instruction to the smart mobile terminal 120 or forwards it to the smart glasses 110 through the smart mobile terminal 120 for execution. At the same time, the cloud server 130 converts the sixth text into corresponding speech and forwards the speech to the smart glasses 110 through the smart mobile terminal 120 for playback.

[0067] While executing the above series of operations, the cloud server 130 can also generate chat records in real time and save them in the historical chat record database for subsequent use (such as use in task requests for recognition operations).

[0068] To reduce the computational burden of a single server and improve data processing efficiency, in another embodiment, the cloud server can be a distributed server cluster consisting of multiple servers. Each of these servers is equipped with a large language model (model server), a historical chat history database (storage server), a speech-to-text engine (speech-to-text server), a text-to-speech engine (text-to-speech server), and a prompt generator (prompt server). The model server interacts with the storage server, speech-to-text server, text-to-speech server, and prompt server to perform the aforementioned operations, such as obtaining historical chat history, converting speech to text, generating prompt information, obtaining the second text, and converting text to speech.

[0069] Optionally, in another embodiment of the present application, as shown in FIG4 , the cloud server 130 includes a model server 131 , and the large language model is configured in the model server 131 .

[0070] The smart mobile terminal 120 is also used to convert the first voice into the first text, perform semantic analysis on the first text and generate the prompt information based on the analyzed semantics, and send the positioning information, the first text, the prompt information and the image to the model server 131.

[0071] The model server 131 is used to obtain the historical chat record, and obtain the second text through the large language model based on the historical chat record, the location information, the first text, the prompt information and the image, and send it to the smart mobile terminal 120.

[0072] The intelligent mobile terminal 120 is further configured to convert the second text into the second voice.

[0073] That is, the smart mobile terminal 120 may also perform the first speech-to-text conversion, prompt information generation, and second text-to-speech conversion operations through the speech-to-text engine, prompt generator, and text-to-speech engine built into the smart mobile terminal 120 .

[0074] Optionally, in another embodiment, the cloud server 130 further includes: a speech-to-text server 132, a prompt server 133, and a text-to-speech server 134, and the smart mobile terminal 120 is further used to: send the first speech to the speech-to-text server 132, so that the first speech is converted into the first text by the speech-to-text server 132; send the first text to the prompt server 133, so that the prompt server 133 performs semantic analysis on the first text and generates the prompt information according to the analyzed semantics; and send the second text to the text-to-speech server 134, so that the second text is converted into the second speech by the text-to-speech server 134.

[0075] That is, the smart mobile terminal 120 can also perform the first speech-to-text conversion, prompt information generation and second text-to-speech conversion operations by interacting with the speech-to-text server 132 , the prompt server 133 and the text-to-speech server 134 .

[0076] Optionally, in another embodiment, the cloud server 130 further includes a storage server 135;

[0077] The smart mobile terminal 120 is further configured to generate real-time chat records and store the real-time chat records in a database on the smart mobile terminal 120 or send the real-time chat records to the storage server 135 for storage;

[0078] The model server 131 is further configured to obtain the historical chat records sent by the smart mobile terminal 120 , or to obtain the historical chat records from the storage server 135 .

[0079] Specifically, the smart mobile terminal 120 or the storage server 135 is configured with a database for storing historical chat records. During human-computer interaction (e.g., chatting) between a user and the large language model, the smart mobile terminal 120 can generate a real-time chat record based on the received or sent data whenever it exchanges data with the smart glasses 110 and the various servers of the cloud server 130. The generated real-time chat record is then stored in the database. The real-time chat record may include, but is not limited to, at least one of received or sent images, voice, and text. Furthermore, the real-time chat record may also include at least one of the time of data exchange, identification information of both parties to the data exchange, and user identification information. Optionally, the real-time chat record may also be generated by the smart glasses 110 or the model server 131 and stored in the database of the storage server 135. Alternatively, each server in the smart glasses 110, the smart mobile terminal 120, and the cloud server 130 may generate its own corresponding real-time chat record and send it to the storage server 135, which then organizes and stores it in the database.

[0080] Optionally, in another embodiment, the cloud server 130 includes: a model server 131, a speech-to-text server 132, a prompt server 133, a text-to-speech server 134 and a storage server 135, and the large language model is configured in the model server 131.

[0081] The model server 131 is also used to generate real-time chat records and send the real-time chat records to the storage server 135 for storage.

[0082] The model server 131 is also used to: obtain the historical chat record from the storage server 135; send the first voice to the voice-to-text server 132, so that the first voice is converted into the first text by the voice-to-text server 132; send the first text to the prompt server 133, so that the prompt server 133 performs semantic analysis on the first text and generates the prompt information according to the analyzed semantics; and send the second text to the text-to-speech server 134, so that the second text is converted into the second voice by the text-to-speech server 134.

[0083] Optionally, in another embodiment, the cloud server 130 includes a model server 131 , and the large language model is configured in the model server 131 .

[0084] The smart glasses 110 are further configured to convert the first speech into a first text, and send the first text and the image to the smart mobile terminal 120 .

[0085] The smart mobile terminal 120 is further configured to perform semantic analysis on the first text and generate the prompt information according to the analyzed semantics, and send the positioning information, the first text, the prompt information and the image to the model server 131 .

[0086] The model server 131 is used to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the location information, the first text, the prompt information and the image, and send it to the smart mobile terminal 120.

[0087] The smart mobile terminal 120 is further configured to send the second text to the smart glasses 110 .

[0088] The smart glasses 110 are further configured to convert the second text into the second voice.

[0089] Furthermore, cloud server 130 also includes a prompt server 133 and a storage server 135. Smart glasses 110 are further configured to send the first text to prompt server 133, which performs semantic parsing on the first text and generates the prompt information based on the parsed semantics. Model server 131 is further configured to retrieve the historical chat records from storage server 135.

[0090] Optionally, in another embodiment, the smart glasses 110 further include a positioning component, and the smart glasses 110 are further configured to obtain the positioning information through the positioning component and send the first voice, the positioning information, and the image to the cloud server 130. In this case, the cloud server 130 is further configured to obtain the second text based on the positioning information, the first text, the prompt information, and the image using the large language model.

[0091] Specifically, the smart glasses 110 can communicate directly with the cloud server 130, that is, while obtaining the first voice, obtain the positioning information through its own positioning component, and send the first voice, the positioning information and the image to the cloud server 130. After receiving the first voice, the positioning information and the image, the cloud server 130 converts the first voice into a first text, performs semantic analysis on the first text and generates prompt information based on the parsed semantics. Through the large language model, based on the positioning information, the first text, the prompt information and the image, the second text is obtained, and the second text is converted into a second voice and returned to the smart glasses 110 for playback. Alternatively, the cloud server 130 can also send the second voice to the smart mobile terminal 120 associated with the smart glasses 110 so that it can be played through the smart mobile terminal 120 or forwarded to the smart glasses 110 through the smart mobile terminal 120.

[0092] Direct communication between the smart glasses 110 and the cloud server 130 can reduce the data loss rate and improve the accuracy and efficiency of data processing.

[0093] Optionally, before positioning, the smart glasses 110 can first determine whether they are equipped with a positioning component based on pre-stored device configuration information. If they have a positioning component, the positioning information is obtained through the positioning component. If they are not equipped with a positioning component, they search for nearby connectable devices via Bluetooth, establish a connection relationship with the searched connectable device, and obtain the positioning information through the connectable device. Alternatively, if they are not equipped with a positioning component, it is determined whether a connection relationship has been established with the associated smart mobile terminal 120. If a connection relationship has been established, the positioning information is obtained from the smart mobile terminal 120. Furthermore, if the smart glasses 110 are not equipped with a positioning component, and there is no connected smart mobile terminal 120, and no connectable device is searched, the smart glasses 110 obtain the last sent positioning information associated with the user from the historical chat record database as the positioning information and mark it so that the cloud server 130 can verify the positioning information based on the mark. Verification method, for example, if the user is bound to a companion, the cloud server 130 can request the companion's smart glasses or smart mobile terminal to obtain positioning information, and verify the marked positioning information based on the obtained positioning information. If the position deviation between the two is less than the preset value, the second text is obtained based on the historical chat record, the positioning information, the first text, the prompt information and the image through the large language model. Otherwise, the second text is obtained through the large language model based on the historical chat record, the first text, the prompt information and the image.

[0094] Optionally, before establishing a data connection with a connectable device, the smart glasses 110 may also use a preset password to verify the connectable device, or only connect to connectable devices in a preset security list to improve the security of data transmission.

[0095] Furthermore, the smart glasses 110 are further configured to transmit the first voice to the smart mobile terminal 120, so that the smart mobile terminal 120 generates the prompt information based on the first voice. The smart glasses 110 are further configured to transmit the prompt information, the first voice, the location information, and the image to the cloud server 130. The cloud server 130 is further configured to obtain historical chat records and, using the large language model, obtain the second text based on the historical chat records, the location information, the first text, the prompt information, and the image.

[0096] Optionally, in another embodiment, the smart glasses 110 further include at least one control button, which includes a physical button and / or a virtual button based on a touch sensor, and the smart glasses 110 are further used to: activate the chat function of the smart glasses 110 in response to an activation instruction, wherein the activation instruction comes from the smart mobile terminal 120 or a virtual assistant program built into the smart glasses 110; capture the image through the camera in response to a photo-taking instruction, and the photo-taking instruction is triggered by the user through the photo-taking button in the control button; output a first prompt tone to prompt the user to ask a question; activate the microphone to start picking up the first voice in response to a listening instruction, and the listening instruction is triggered based on an event in which the user presses the chat button in the control button; stop picking up the first voice in response to an end instruction, and output a second prompt tone to prompt the user to end picking up the voice, and the end instruction is triggered based on an event in which the user releases the chat button or the microphone is idle for more than a preset time.

[0097] Optionally, while sending the first voice or first text, the smart glasses 110 or the smart mobile terminal 120 may also send one or more sensory data obtained by built-in sensors, including pedometer step data, user posture, sensor data such as IMU (inertial measurement unit) data, electronic compass direction data, touch sensor data, and environmental data such as temperature or humidity, to the large language model so that the large language model can better respond to user queries. For example, if a user asks, "How long will it take me to get to the place in the picture?" the large language model will combine the location of the recognized place in the image and the pedometer step data to correctly answer the user's question about the remaining time, such as, "You still need to walk for half an hour to get to the place."

[0098] In this embodiment, by utilizing a large language model, smart glasses enable human-computer task interaction based on captured images, thereby enriching their functionality. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses can be further enhanced. Furthermore, because the smart glasses can utilize nearby connected devices for positioning and data forwarding even when they lack their own wireless network or positioning components, the success rate of data processing can be improved.

[0099] Referring to Figure 5 , Figure 5 is a schematic diagram of the internal structure of smart glasses provided in accordance with an embodiment of the present application. Figure 6 is a schematic diagram of the external structure of smart glasses provided in accordance with an embodiment of the present application. For ease of illustration, only the portions relevant to the present embodiment are shown in the figures. As shown in Figures 5 and 6 , smart glasses 200 include: a frame 201, at least one temple 202, at least one microphone 203, at least one speaker 204, at least one processor 205, at least one memory 206, and at least one sensor 207. The at least one sensor 207 includes a camera 2071. The camera 2071 can be located on the left side of the frame 201 (as shown in Figure 6 ), on the right side of the frame 201, or in the center of the frame 201, without specific limitation in this application. Figures 5 and 6 are merely preferred examples. In actual applications, the smart glasses 200 may have fewer or more components than those shown in Figures 5 and 6 .

[0100] The frame 201 may be, for example, a front frame with a lens (eg, a sunglass lens, a transparent lens, or a corrective lens). The at least one temple 202 may include, for example, a left temple and a right temple.

[0101] The temple 202 is connected to the frame 201, and the processor 205 is electrically connected to the microphone 203, the speaker 204, the memory 206, and the sensor 207. The microphone 203, the speaker 204, the processor 205, the memory 206, and the sensor 207 are disposed on at least one of the temples 202 and / or the frame 201. Preferably, the temple 202 is detachably connected to the frame 201.

[0102] The processor 205 includes a central processing unit (CPU) and a digital signal processor (DSP). The DSP is used to process the voice data acquired by the microphone 203. The CPU is preferably a microcontroller unit (MCU).

[0103] The memory 206 is a non-transitory memory, which may specifically include: RAM (Random Access Memory) and flash memory components, storing one or more programs executable by the processor 205, wherein the one or more programs include multiple instructions. The multiple instructions are used to: obtain sensor data through at least one sensor 207, wherein at least one sensor 207 includes a camera 2071, and the sensor data includes images captured by the camera 2071; obtain a first voice of a user through the microphone, wherein the first voice includes a question posed by the user; obtain a second voice including an answer to the question based on the first voice and the sensor data using a large language model, wherein the large language model is configured in the smart glasses, smart mobile terminal, or cloud server; and play the second voice through the speaker 204.

[0104] The large language model may be a GAILLM model, an MLLM model, or other models with similar functions.

[0105] Optionally, in other embodiments of the present application, the smart glasses 200 further include at least one control button 208 electrically connected to the processor 205, the control button 208 including a physical button and / or a virtual button based on a touch sensor, and the multiple instructions are further used to: activate the chat function of the smart glasses 200 in response to an activation instruction, wherein the activation instruction comes from an associated smart mobile terminal or a virtual assistant program built into the smart glasses 200; capture the image through the camera 2071 in response to a photo-taking instruction, wherein the photo-taking instruction is triggered by the user through the photo-taking button in the control button 208; output a first A prompt sound is output to prompt the user to ask a question; in response to a listening instruction, the microphone 203 is activated to start picking up the first voice, wherein the listening instruction is triggered based on the event that the user presses the chat button in the control button 208; in response to an end instruction, the first voice is stopped from being picked up, a second prompt sound is output to prompt the user to end the sound pickup, and the operation of obtaining the second voice containing the answer to the question based on the first voice and the sensor data through the large language model is executed, wherein the end instruction is triggered based on the event that the user releases the chat button or the event that the idle time of the microphone 203 exceeds the preset time.

[0106] Optionally, in other embodiments of the present application, the at least one sensor 207 further includes a position sensor 2072 (not shown in Figures 5 and 6), and the sensor data further includes positioning data of the smart glasses 200. The multiple instructions are further configured to obtain the position information of the smart glasses 200 as the positioning data via the position sensor 2072. The function of the position sensor 2072 is similar to the positioning component or satellite positioning system in the control system embodiments shown in Figures 1 and 4, and may specifically be, but is not limited to, a positioning component based on GPS or Beidou satellites.

[0107] Optionally, in other embodiments of the present application, the multiple instructions are also used to: obtain the user's historical chat records, and obtain the second voice based on the first voice, the sensor data and the historical chat records through the large language model.

[0108] Optionally, in other embodiments of the present application, the smart glasses 200 also include a wireless communication component 208 electrically connected to the processor 205, and the multiple instructions are also used to: generate real-time chat records and save them in the memory, or send the real-time chat records to the storage server for storage through the wireless communication component; obtain the historical chat records from the memory or the storage server.

[0109] Specifically, the wireless communication component 208 includes a wireless signal transceiver and its surrounding circuitry, and may be disposed within the inner cavity of the frame 201 and / or at least one of the temples 202. The wireless signal transceiver may utilize, but is not limited to, at least one of the following protocols for data transmission: WiFi (Wireless Fidelity), NFC (Near Field Communication), ZigBee, UWB (Ultra Wideband), RFID (Radio Frequency Identification), and cellular mobile communication protocols (e.g., 3G / 4G / 5G).

[0110] Optionally, in other embodiments of the present application, the smart glasses 200 also include a Bluetooth component 209 electrically connected to the processor 205, and the multiple instructions are also used to: obtain the positioning data from the smart mobile terminal through the Bluetooth component 209; send the first voice to the voice-to-text server through the wireless communication component 208, so that the first voice is converted into a first text through the voice-to-text server; send the first text to the prompt server through the wireless communication component 208, so that the prompt server performs semantic analysis on the first text and generates the prompt information based on the analyzed semantics; obtain a second text containing an answer to the question based on the prompt information, the first text, the sensor data and the historical chat record through the large language model; and send the second text to the text-to-speech server through the wireless communication component 208, so that the second text is converted into the second voice through the text-to-speech server.

[0111] Optionally, in other embodiments of the present application, the multiple instructions are also used to: send the prompt information, the first text, the sensor data and the historical chat record to the smart mobile terminal through the Bluetooth component 209, so as to obtain the second text through the large language model on the smart mobile terminal, wherein the second text contains the answer to the question; or, send the prompt information, the first text, the sensor data and the historical chat record to the model server through the wireless communication component 208, so as to obtain the second text through the large language model on the model server.

[0112] Optionally, in other embodiments of the present application, the smart glasses 200 further include other input devices electrically connected to the processor 205, such as a power button, a touch screen, etc. Furthermore, the touch screen can also serve as an output device for displaying the real-time chat records or user-specified historical chat records in the other embodiments described above.

[0113] Optionally, in other embodiments of the present application, the smart glasses 200 further include an indicator light and / or a buzzer electrically connected to the processor 205. The multiple instructions are further configured to output prompt information via the indicator light and / or the buzzer. The prompt information is configured to indicate the status of the smart glasses 200, including an operating state and an idle state. The operating state includes: a state in which voice pickup has started, a state in which voice pickup has ended, and a state in which voice processing has been performed. The indicator light may optionally be an LED (Light Emitting Diode).

[0114] Specifically, in addition to using a beep notification to indicate that the smart glasses are listening to the user's voice, to indicate that the voice has been sent to the large language model and is waiting for a response, the smart glasses can also use an LED to indicate whether the smart glasses are listening to the user's voice (for example, green), waiting for a response from the large language model (red), or completely idle (off). And after playing the voice obtained by the large language model (such as the second voice), if the chat function of the smart glasses is turned off because the microphone has been idle for more than a preset period of time, the LED will also be turned off. At this time, if the user sees that this LED is off, they know that they can use the method of pressing and holding the virtual button or the wake-up word to reactivate the smart glasses to listen to the user's voice again.

[0115] Optionally, in other embodiments of the present application, the at least one sensor 207 further includes: at least one of an inertial measurement unit (IMU) sensor, a temperature sensor, a proximity sensor, a humidity sensor, an electronic compass, a timer, and a pedometer.

[0116] The plurality of instructions are further configured to, in response to the model server's sensory data acquisition request, acquire sensory data from the at least one device and transmit it to the model server, so that the model server, through the large language model, obtains the second text based on the historical chat history, the location information, the first text, the prompt information, the image, and the sensory data. It is understood that the large language model may also output other text containing the sensory data acquisition request. The model server generates the sensory data acquisition request based on the other text and transmits it to the smart glasses 200, so that the smart glasses 200 acquire the sensory data indicated by the request according to the sensory data acquisition request. For example, if a user submits a task request such as "How long do I need to walk to reach the location in the picture?", the large language model will output text containing a pedometer data acquisition instruction. The model server, based on the pedometer data acquisition instruction, will transmit a pedometer data acquisition request to the smart glasses 200 or a smart mobile terminal associated with the smart glasses 200, to acquire pedometer data from the smart glasses 200 or a smart mobile terminal associated with the smart glasses 200 and input it into the large language model. The large language model calculates the time it takes for the user to arrive at the location based on the step count data of the pedometer and the positions of the objects in the recognized image, and outputs a second text containing the information "You still need to walk for half an hour to reach the location."

[0117] Optionally, in other embodiments of the present application, the smart glasses 200 further include a voice biometrics module, which is configured to identify the user of the smart glasses 200 using the acquired voiceprint of the user, thereby enabling the aforementioned voice control function based on the smart glasses 200. Since only the owner of the smart glasses can activate the function provided by the smart glasses using the large language model, the security of the device can be improved.

[0118] Furthermore, the smart glasses 200 also include a battery 210 for providing power support for the electronic devices of the smart glasses 200, such as the microphone 203, the speaker 204, the processor 205, the memory 206, the at least one sensor 207 and other components.

[0119] The electronic components of the smart glasses can be connected via a bus.

[0120] It should be noted that the relationships between the components of the aforementioned smart glasses can be either replacement or superposition. That is, a pair of smart glasses can be equipped with all of the components described in this embodiment, or, depending on the needs, selectively install a portion of the components. In the case of replacement, the smart glasses also include a peripheral connection interface, which can be, for example, at least one of a PS / 2 interface, a serial interface, a parallel interface, an IEEE1394 interface, and a USB (Universal Serial Bus) interface. The functionality of the replaced component can be implemented by a peripheral connected to the interface, such as an external speaker or sensor.

[0121] For details not covered in this embodiment, please refer to the relevant descriptions in the embodiments shown in Figures 1 to 4 and Figures 7 to 12 below, and will not be repeated here.

[0122] In this embodiment, by utilizing a large language model, smart glasses enable human-computer task interaction based on captured images, thereby enriching their functionality. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses can be further enhanced. Furthermore, since the smart glasses can also utilize nearby connected devices for positioning and data forwarding when they lack their own wireless networks or positioning components, the success rate of data processing can be improved.

[0123] See Figure 7, which is a flow chart of an implementation method of a smart wearable device control method based on a large language model provided by an embodiment of the present application. The control method can be applied to a smart mobile terminal, such as the smart mobile terminal 120 in the embodiments shown in Figures 1 and 4. The smart wearable device can implement the control method by means of data interaction with the smart mobile terminal 120. Among them, the smart wearable device can include, but is not limited to: a smart helmet, a smart headset, a smart earring, a smart watch, and smart glasses as shown in Figures 5 and 6. As shown in Figure 7, the control method includes the following steps:

[0124] S701: Receiving a first voice and an image sent by a smart wearable device via Bluetooth, wherein the first voice includes a question raised by a user;

[0125] S702: Convert the first speech into a first text, perform semantic analysis on the first text, and generate prompt information according to the analyzed semantics;

[0126] S703: Obtaining a second text based on the image, the first text, and the prompt information using a large language model, wherein the second text includes an answer to the question, and the large language model is configured on the smart mobile terminal or the cloud server;

[0127] S704: Convert the second text into a second voice, and send the second voice to the smart wearable device via Bluetooth for playback.

[0128] Optionally, in other embodiments of the present application, the method also includes: obtaining location information and / or the user's historical chat records; and obtaining the second text based on the image, the first text, the prompt information, and the location information and / or the historical chat records through a large language model.

[0129] Optionally, in other embodiments of the present application, converting the first speech into the first text may further specifically include: sending the first speech to a speech-to-text server via a wireless network, so that the speech-to-text server converts the first speech into the first text.

[0130] Performing semantic parsing on the first text and generating prompt information according to the parsed semantics may specifically include: sending the first text to a prompt server via the wireless network, so that the prompt server performs semantic parsing on the first text and generates the prompt information according to the parsed semantics.

[0131] Converting the second text into a second voice may further include: sending the second text to a text-to-speech server via the wireless network, so that the text-to-speech server converts the second text into the second voice.

[0132] Optionally, in other embodiments of the present application, after converting the first speech into the first text, the method further includes: generating a real-time chat record based on the first text and the image, and storing the real-time chat record in a database or storage server of the smart mobile terminal; displaying the real-time chat record on a display screen; after obtaining the second text, associating the second text with the real-time chat record in the database or the storage server; and displaying the reply in the second text on the display screen. The display screen may be a built-in or external display screen of the smart mobile terminal.

[0133] Acquiring the historical chat records of the user may further include: acquiring the historical chat records during the current chat from the database or the storage server.

[0134] Optionally, in other embodiments of the present application, the prompt information is used to prompt the large language model to perform at least one of the following tasks based on the image: image recognition, information sharing, making a call, navigation, searching the Internet, calling a network server, and translation.

[0135] The following will further illustrate the above-mentioned smart glasses control system and smart wearable device control method with reference to multiple application examples.

[0136] Application Example 1

[0137] The large language model is configured in server 130. Smart glasses 110 are responsible for voice recording, playback, and taking photos. The smartphone or smartwatch 120 is responsible for GPS positioning. The cloud server is responsible for voice-to-text conversion, text-to-speech conversion, accessing chat history, generating prompts, and obtaining answers to user questions using the large language model. Prompts can also be generated by the smartphone or smartwatch 120.

[0138] As shown in FIG8 , after the user starts the chat function through a mobile application (APP) or a virtual assistant program (such as Siri / OK Google, etc.) running on a smartphone or smart watch 120, if the user presses and holds the touch sensor-based virtual button on the temple of the smart glasses 110, the smart glasses 110 call the built-in camera to take a picture, and at the same time, outputs a prompt sound through the built-in speaker to prompt the user that the microphone on the smart glasses 110 is ready to listen to the user's speech.

[0139] While holding down the virtual button, the user asks a question, such as "Where in Hong Kong is this photo?" The smart glasses 110 pick up the first voice message containing the question through the microphone. When the user releases the virtual button, the smart glasses 110 transmit the first voice message and the captured photo to the mobile app on the smartphone or smartwatch 120 via Bluetooth and output another notification sound to inform the user that the task request has been sent.

[0140] The mobile app then obtains GPS positioning information and sends the first voice, the photo, and the GPS positioning information to the cloud server 130. The cloud server 130 converts the first voice into a first text, performs semantic analysis on the first text, generates prompt information based on the parsed semantics, and obtains the user's historical chat history. This historical chat history can, for example, be the text corresponding to the user's voice and / or the photo sent during the user's most recent chat with the large language model after the chat function was activated and before the current time point. It is understandable that a question and answer between the user and the large language model can be considered a chat.

[0141] Afterwards, the cloud server 130 inputs the captured photo, the first text, the prompt information, the GPS location information, and the historical chat records into the large language model, obtains a second text output by the large language model containing the answer to the user's question, converts the second text into a second voice, and sends the second voice to the smartphone or smartwatch 120. The smartphone or smartwatch 120 then sends the received second voice to the smart glasses 110 via Bluetooth, so that the smart glasses 110 can play it using the built-in speakers.

[0142] The user can then continue to ask the next question until the user stops the chat function through the mobile app.

[0143] Furthermore, the user can set the language he or she uses through the mobile APP, or the user language can be automatically detected by the speech-to-text engine, and the prompt information generated by the cloud server 130 also includes the user language information.

[0144] Furthermore, the user can set the playback speed of the second voice through the mobile APP, for example, setting it to normal, 1.25 times faster or slower, 1.5 times faster or slower, or 2 times faster or slower.

[0145] Furthermore, the smartphone or smartwatch 120 can generate real-time chat records based on the received data during the chat function activation period, and display the real-time chat records to the user through the mobile app, while also associating the real-time chat records with the user account and storing them in a historical chat record database. The historical chat record database can be configured in a storage server or in the smartphone or smartwatch 120.

[0146] Furthermore, based on the user's operation on the mobile APP, the smartphone or smartwatch 120 can also export the chat history directed by the user's operation as a file to a social media platform, such as Facebook.

[0147] Furthermore, users can view complete chat records or chat records of specified dates or content through the mobile app.

[0148] Furthermore, when the smart glasses 110 play the second voice, the user can slide the temples of the smart glasses 110 toward the ears to increase the volume, or slide the temples of the smart glasses 110 away from the ears to decrease the volume).

[0149] Application Example 2

[0150] The large language model is configured in the server 130. As shown in FIG9 , unlike Application Example 1, in Application Example 2, when the user stops speaking and remains idle for a period of time (i.e., the microphone is idle for longer than a preset time), the smart glasses 110 stop picking up the sound and send the first voice and the captured photo to the smartphone or smartwatch 120.

[0151] Application Example 3

[0152] The large language model is configured in server 130. As shown in Figure 10, unlike Application Example 1, in Application Example 3, smart glasses 110 are responsible for voice capture, voice playback, photo taking, voice-to-text conversion, and text-to-speech conversion. Smartphone or smartwatch 120 is responsible for GPS positioning. Cloud server 130 is responsible for generating prompt information, obtaining historical chat records, and obtaining answers to user questions through the large language model.

[0153] Application Example 4

[0154] The large language model is configured in server 130. As shown in Figure 11, unlike Application Example 1, in Application Example 4, smart glasses 110 are responsible for voice recording, voice playback, and taking photos, while smartphone or smartwatch 120 is responsible for GPS positioning, voice-to-text conversion, and text-to-speech conversion. Cloud server 130 is responsible for generating prompts, obtaining historical chat logs, and obtaining answers to user questions through the large language model.

[0155] Application Example 5

[0156] The large language model is configured in server 130. As shown in Figure 12, unlike Application Example 1, in Application Example 5, smart glasses 120 communicate directly with cloud server 130. Smart glasses 110 are responsible for voice recording, voice playback, GPS positioning, and taking photos, while cloud server 130 is responsible for voice-to-text conversion, text-to-speech conversion, generating prompts, obtaining historical chat logs, and obtaining answers to user questions through the large language model.

[0157] For details not covered in this embodiment, please refer to the relevant descriptions in the embodiments shown in Figures 1 to 6 above, and will not be repeated here.

[0158] In this embodiment, by utilizing a large language model, human-computer task interaction based on captured images is implemented in smart glasses, thereby enriching the functionality of the smart glasses. Furthermore, due to the scalability and self-creativity of the large language model, the intelligence and interactivity of the smart glasses can be further improved.

[0159] The embodiments of the present application also provide a non-transitory computer-readable storage medium, which may be provided in the smart glasses or smart wearable devices in the above embodiments. The non-transitory computer-readable storage medium may be the memory 206 in the embodiment shown in FIG5 . The computer-readable storage medium stores a computer program, which, when executed by the processor, implements the smart wearable device control method based on the large language model described in the above embodiments. Furthermore, the computer-storable medium may also be various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk, or an optical disk.

[0160] In the several embodiments provided herein, it should be understood that the disclosed smart glasses, control systems, and smart wearable device control methods can be implemented in other ways. For example, multiple modules or components can be combined or integrated into another system, or some features can be omitted or not implemented. In addition, the connections or direct connections or communication connections shown or discussed can be indirect connections or communication connections through some interfaces, devices, or modules, and can be electrical, mechanical, or other forms.

[0161] It should be noted that for the aforementioned embodiments of the smart wearable device control method, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited to the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0162] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0163] The above is a description of the smart glasses, control system, and smart wearable device control method provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. An intelligent glasses control system based on a large language model, characterized in that, The system includes: smart glasses, a smart mobile terminal, and a cloud server, where the smart glasses include a microphone, a speaker, a camera, and a Bluetooth component, and the large language model is configured in the cloud server; The smart glasses are used to capture an image through the camera, obtain the user's first voice through the microphone, and send the image and the first voice to the smart mobile terminal through the Bluetooth component, where the first voice contains the question raised by the user; The smart mobile terminal is used to send the first voice and the image to the cloud server; The cloud server is used to: convert the first voice into a first text; perform semantic parsing on the first text and generate a prompt message according to the parsed semantics; through the large language model, obtain a second text according to the first text, the prompt message, and the image, where the second text contains the answer to the question; convert the second text into a second voice; and send the second voice to the smart mobile terminal; The smart mobile terminal is further used to send the second voice to the smart glasses; The smart glasses are further used to receive the second voice through the Bluetooth component and play the second voice through the speaker.

2. The smart glasses control system according to claim 1, wherein The smart mobile terminal is further used to obtain location information and send the location information, the first voice, and the image to the cloud server, where the location information is obtained by the smart mobile terminal through a built-in satellite positioning system or is obtained by the smart glasses and sent to the smart mobile terminal; The cloud server is further used to obtain the second text through the large language model according to the location information, the first text, the prompt message, and the image.

3. The smart glasses control system according to claim 1, wherein The cloud server is further used to obtain historical chat records and obtain the second text through the large language model according to the historical chat records, the first text, the prompt message, and the image.

4. The smart glasses control system according to claim 3, wherein The smart mobile terminal is further used to obtain location information and send the location information, the first voice, and the image to the cloud server, where the location information is obtained by the smart mobile terminal through a built-in satellite positioning system or is obtained by the smart glasses and sent to the smart mobile terminal; The cloud server is further used to obtain the historical chat records and obtain the second text through the large language model according to the historical chat records, the location information, the first text, the prompt message, and the image.

5. The intelligent glasses control system according to claim 4, wherein The cloud server includes a model server, and the large language model is configured in the model server; The intelligent mobile terminal is further configured to convert the first voice into the first text, perform semantic analysis on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server; The model server is configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send it to the intelligent mobile terminal; The intelligent mobile terminal is further configured to convert the second text into the second voice.

6. The intelligent glasses control system according to claim 5, wherein The cloud server further includes: a voice-to-text server, a prompt server, and a text-to-voice server. The intelligent mobile terminal is further configured to: Send the first voice to the voice-to-text server to convert the first voice into the first text through the voice-to-text server; Send the first text to the prompt server to perform semantic analysis on the first text and generate the prompt information according to the parsed semantics; and Send the second text to the text-to-voice server to convert the second text into the second voice through the text-to-voice server.

7. The intelligent glasses control system according to claim 6, characterized in that, The cloud server further includes a storage server; The intelligent mobile terminal is further configured to generate a real-time chat record, and save the real-time chat record in a database on the intelligent mobile terminal, or send it to the storage server for saving; The model server is further configured to obtain the historical chat record sent by the intelligent mobile terminal, or obtain the historical chat record from the storage server.

8. The intelligent glasses control system according to claim 4, wherein The cloud server further includes: a model server, a voice-to-text server, a prompt server, a text-to-voice server, and a storage server. The large language model is configured in the model server; The model server is further configured to generate a real-time chat record and send the real-time chat record to the storage server for saving; The model server is further configured to: Obtain the historical chat record from the storage server; Send the first voice to the voice-to-text server to convert the first voice into the first text through the voice-to-text server; Send the first text to the prompt server to perform semantic analysis on the first text and generate the prompt information according to the parsed semantics; and Send the second text to the text-to-voice server to convert the second text into the second voice through the text-to-voice server.

9. The intelligent glasses control system according to claim 4, wherein, wherein, The cloud server includes a model server, and the large language model is configured in the model server; The intelligent glasses are further configured to convert the first voice into the first text, and send the first text and the image to the intelligent mobile terminal; The intelligent mobile terminal is further configured to perform semantic parsing on the first text, generate the prompt information according to the parsed semantics, and send the positioning information, the first text, the prompt information, and the image to the model server; The model server is configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image, and send it to the intelligent mobile terminal; The intelligent mobile terminal is further configured to send the second text to the intelligent glasses; The intelligent glasses are further configured to convert the second text into the second voice.

10. The intelligent glasses control system according to claim 9, characterized in that, characterized in that, The cloud server further includes: a prompt server and a storage server; The intelligent glasses are further configured to send the first text to the prompt server, so that the prompt server performs semantic parsing on the first text and generates the prompt information according to the parsed semantics; The model server is further configured to obtain the historical chat record from the storage server.

11. The intelligent glasses control system according to claim 1, wherein The intelligent glasses further include a positioning component, and the intelligent glasses are further configured to: obtain positioning information through the positioning component, and send the first voice, the positioning information, and the image to the cloud server; The cloud server is further configured to obtain the second text according to the positioning information, the first text, the prompt information, and the image through the large language model.

12. The intelligent glasses control system according to claim 11, wherein The intelligent glasses are further configured to send the first voice to the intelligent mobile terminal, so that the intelligent mobile terminal generates the prompt information according to the first voice; The intelligent glasses are further configured to send the prompt information, the first voice, the positioning information, and the image to the cloud server; The cloud server is further configured to obtain the historical chat record, and through the large language model, obtain the second text according to the historical chat record, the positioning information, the first text, the prompt information, and the image.

13. The intelligent glasses control system according to claim 1, wherein The intelligent glasses further include at least one control button, the control button includes a physical button and / or a virtual button based on a touch sensor, and the intelligent glasses are further configured to: Respond to an activation instruction to activate the chat function of the intelligent glasses, where the activation instruction comes from the intelligent mobile terminal or a virtual assistant program built in the intelligent glasses; Respond to a photographing instruction to capture the image through the camera, and the photographing instruction is triggered by the user through the photographing button in the control buttons; Output a first prompt sound to prompt the user to ask a question; Respond to a listening instruction to activate the microphone to start picking up the first voice, and the listening instruction is triggered based on an event that the user presses the chat button in the control buttons; In response to an end instruction, stop picking up the first voice, and output a second prompt tone to prompt the user to end voice pickup. The end instruction is triggered based on an event that the user releases the chat button or an event that the idle time of the microphone exceeds a preset time.

14. The intelligent glasses control system according to claim 1, characterized in that, The large language model includes a generative artificial intelligence large language model or a multimodal large language model.

15. An intelligent glasses based on large language models, characterized in that, Comprising: A spectacle frame, at least one temple, a microphone, a speaker, at least one sensor, a processor, and a memory; The at least one temple is connected to the spectacle frame, and the processor is electrically connected to the microphone, the speaker, the at least one sensor, and the memory; One or more programs executable by the processor are stored in the memory. The one or more programs include a plurality of instructions, and the plurality of instructions are used for: Obtain sensing data through the at least one sensor. Wherein, the at least one sensor includes a camera, and the sensing data includes an image captured by the camera; Obtain the first voice of the user through the microphone. The first voice contains a question raised by the user; Through a large language model, obtain a second voice containing a reply to the question according to the first voice and the sensing data. The large language model is configured in the smart glasses or a smart mobile terminal or a cloud server; Play the second voice through the speaker.

16. The smart glasses according to claim 15, characterized in that, The smart glasses further include at least one control button electrically connected to the processor. The control button includes a physical button and / or a virtual button based on a touch sensor. The plurality of instructions are further used for: In response to an activation instruction, activate the chat function of the smart glasses. The activation instruction comes from the smart mobile terminal or a virtual assistant program built in the smart glasses; In response to a photographing instruction, capture the image through the camera. The photographing instruction is triggered by the user through the photographing button in the control button; Output a first prompt tone to prompt the user to ask a question; In response to a listening instruction, activate the microphone to start picking up the first voice. The listening instruction is triggered based on an event that the user presses the chat button in the control button; In response to an end instruction, stop picking up the first voice, output a second prompt tone to prompt the user to end voice pickup, and perform the operation of obtaining a second voice containing a reply to the question according to the first voice and the sensing data through the large language model. The end instruction is triggered based on an event that the user releases the chat button or an event that the idle time of the microphone exceeds a preset time.

17. The smart glasses according to claim 16, characterized in that, The at least one sensor further includes a position sensor, and the sensing data further includes positioning data of the smart glasses. The plurality of instructions are further used for: Obtain the position information of the smart glasses as the positioning data through the position sensor.

18. The smart glasses according to claim 17, characterized in that, The plurality of instructions are further used for: Obtain the user's historical chat record, and through the large language model, obtain the second voice according to the first voice, the sensing data, and the historical chat record.

19. The smart glasses according to claim 18, characterized in that, The smart glasses further include a wireless communication component electrically connected to the processor, and the plurality of instructions are further configured to: Generate a real-time chat record and save it in the memory, or send the real-time chat record to a storage server for saving through the wireless communication component; Obtain the historical chat record from the memory or the storage server.

20. The smart glasses according to claim 19, characterized in that, The smart glasses further include a Bluetooth component electrically connected to the processor, and the plurality of instructions are further configured to: Obtain the positioning data from the smart mobile terminal through the Bluetooth component; Send the first voice to a speech-to-text server through the wireless communication component, so that the speech-to-text server converts the first voice into a first text; Send the first text to a prompt server through the wireless communication component, so that the prompt server performs semantic parsing on the first text and generates the prompt information according to the parsed semantics; Obtain a second text containing a reply to the question according to the prompt information, the first text, the sensing data, and the historical chat record through the large language model; And Send the second text to a text-to-speech server through the wireless communication component, so that the text-to-speech server converts the second text into the second voice.

21. The smart glasses according to claim 20, wherein, The plurality of instructions are further configured to: Send the prompt information, the first text, the sensing data, and the historical chat record to the smart mobile terminal through the Bluetooth component, so as to obtain a second text containing a reply to the question through the large language model on the smart mobile terminal; Or, Send the prompt information, the first text, the sensing data, and the historical chat record to a model server through the wireless communication component, so as to obtain the second text through the large language model on the model server.

22. A control method for an intelligent wearable device based on a large language model, characterized in that, Applied to a smart mobile terminal, the method includes: Receiving a first voice and an image sent by a smart wearable device through Bluetooth, where the first voice includes a question raised by a user; Converting the first voice into a first text, performing semantic parsing on the first text, and generating prompt information according to the parsed semantics; Obtaining a second text containing a reply to the question according to the image, the first text, and the prompt information through a large language model, where the large language model is configured on the smart mobile terminal or a cloud server; Converting the second text into a second voice, and sending the second voice to the smart wearable device through the Bluetooth for playing.

23. The control method according to claim 22, wherein The method further includes: Obtaining positioning information and / or the user's historical chat record; Obtaining the second text according to the image, the first text, the prompt information, and the positioning information and / or the historical chat record through a large language model.

24. The control method according to claim 23, wherein The converting the first voice into a first text includes: Send the first voice to a speech-to-text server via a wireless network to convert the first voice into the first text by the speech-to-text server; The semantic parsing of the first text and generating a prompt message according to the parsed semantics includes: Send the first text to a prompt server via the wireless network to perform semantic parsing on the first text by the prompt server and generate the prompt message according to the parsed semantics; The conversion of the second text into a second voice includes: Send the second text to a text-to-speech server via the wireless network to convert the second text into the second voice by the text-to-speech server.

25. The control method according to claim 23, wherein After converting the first voice into the first text, the method further includes: Generate a real-time chat record according to the first text and the image, and store the real-time chat record in a database or a storage server of the intelligent mobile terminal; Display the real-time chat record through a display screen; After obtaining the second text, associate the second text with the real-time chat record in the database or the storage server; Display the reply in the second text through the display screen; The obtaining of the user's historical chat record includes: Obtain the historical chat record during the current chat from the database or the storage server.

26. The control method according to claim 23, characterized in that The prompt message is used to prompt the large language model to perform at least one of the following tasks based on the image: image recognition, information sharing, making a phone call, navigation, searching the Internet, calling a network server, and translation.

Citation Information

Patent Citations

  • Intelligent wearable terminal, cloud server and data processing method

    CN110287830A

  • Voice information processing method and system for AR glasses

    CN113763940A

  • Intelligent glasses and control method and system thereof

    CN115695620A

  • Multi-model cooperation method based on large-scale language model

    CN116976306A

  • Intelligent glasses, system and control method based on generative artificial intelligence large language model

    CN119002054A